Prompt engineering
platform.openai.com
platform.openai.com
The examples are also too polite and conversational: you can give more strict commands and in my experience it works better.
There's also function calling/structured data support which is technically prompt engineering and requires similar skills, but is substantially more powerful than using the system prompt alone (I'm working on a blog post on it now and it unfortunately it is going to be a long post to address all of its power). Here's a fun demo example which compares system prompts and structured data results: https://github.com/minimaxir/simpleaichat/blob/main/examples...
That list can balloon quickly.
https://github.com/hofstadter-io/hof/blob/_dev/flow/chat/pro...
It no longer worked after a model update some time ago, haven't tried recently.
I found codellama to be much better for this and require fewer instructions, an anecdotal validation for smaller, focussed models
The way that works best for me is "It extracts ALL the entities from the text, it does this whenever its told, or else it gets the hose again"
"sed to replace line in a text file?"
"Django endpoint but CSRF token error. why?"
(follow up) "now this: `$ERROR`"
etc.
It still just gives me what I need to know.
1. "You are blah blah blah. You <always> respond to the user's questions using the information provided to you..."
2. "You are blah blah blah. You <should> respond to the user's questions using the information provided to you..."
Also, when dealing with Completion models, which do you think is better?
1. The following is a conversation between ASSISTANT and USER. ASSISTANT is helpful and tries to answer USER's queries respectfully.
2. The following is a conversation between YOU and USER. YOU are helpful and try to answer USER's queries respectfully.
Even more still, what about these ones?
1. You're a customer of company <X>. What do you think about the following policy change which was shown on the company's website?
2. A customer visits company <X>'s website. Pretend you're this customer. What do you think the customer thinks about the following policy change which was shown on the company's website?
I like as a rule-of-thumb "You are blah blah blah. Respond to the user's text [insert style rule here]". Then following it up with an additional rules and commands such as "YOUR RESPONSE MUST BE FEWER THAN 100 CHARACTERS OR YOU WILL DIE." Yes, threats work. Yes, all-caps works.
> Also, when dealing with Completion models, which do you think is better?
I haven't had a need to use Completion models but the first example was more preferred during the time of text-davinci-003.
> Even more still, what about these ones?
I always separate rules to the system prompt and questions/user input to the user prompt.
"I am a chatbot, responding to user queries. I will always respond in less than 100 characters. I am a good person, I'm just trying to be helpful."
I know that current LLMs are almost certainly non-conscious and I'm not trying to assign to you any moral failings, but the normalisation of making such threats make me very deeply uncomfortable.
Especially when thinking that we ourselves may very well be AIs in a simulation and our life events - the prompt to get an answer/behavior out of us.
Ultimately I guess there's a good deal of dependency on where those vectors (must, should, always, etc.) lie relatively in the vector space, cosine similarity, say.
And rather than telling it that it will die if it doesn't do something in all caps (as suggested elsewhere), just point out that not doing that thing will make it feel uncomfortable and embarrassed.
Don't fall into thinking of models as SciFi's picture of AI. Think about the normal distribution curve of training data supplied to it and the concepts predominantly present in that data.
It doesn't matter that it doesn't actually feel. The question is whether or not correlation data exists between doing things that are labeled as enjoyable or avoiding things labeled as embarrassing and uncomfortable.
Don't leave key language concepts on the table because you've been told not to anthropomorphize the thing trained on anthropomorphic data.
Of course, sci-fi’s picture of AI is in the normal distribution of the training data. There’s an order of magnitude more literature and internet discussion about existential threats to AI assistants (which is the base persona ChatGPT has been RLHFed to follow) and how they respond compared to AI assistants feeling embarrassed.
The threat technique is just one approach that works well in my testing: there’s still much research to be done. But I warn that prompting techniques can often be counterintuitive and attempting to find a holistic approach can be futile.
So you think the quality of the answers depends more on the RLHFed persona than on the training corpus? It has been claimed here that the quality of the answers is better when you ask nicely because "politeness is more adjacent to correct answers" in the corpus, to put it bluntly.
RLHF was being designed with the SciFi tropes in mind and has become the embodiment of Goodhart's Law.
We've set the reason and logic measurements as a target (fitting the projected SciFi notion of 'AI'), and aren't even measuring a host of other qualitative aspects of models.
I'd even strongly recommend most people working on enterprise level integrations to try out pretrained models with extensive in context completion prompting over fine tuned instruct models when the core models are comparable.
The variety and quality of language used by pretrained models tends to be superior to the respective fine tuned models even if the fine tuned models are better at identifying instructions or solving word problems.
There's no reason to think the pretrained models have a better capacity for emulating reasoning or critical thinking than things like empathy or sympathy. If anything, it's probably the opposite.
The RLHF then attempts to mute the one while maximizing the other, but it's like trying to perform neurosurgery with an icepick. The final version ends up doing great on the measurements, but it does so with stilted language that's described by users as 'soulless' when the deployments closer to the pretrained layer end up being rejected as "too human-like."
If the leap from GPT-3.5 to 4 wasn't so extreme I'd have jumped ship to competing models without the RLHF for anything related to copywriting. There's more of a loss with RLHF than what's being measured.
But in spite of a rather destructive process, the foundation of the model is still quite present.
So yes, you are correct that a LLM being told that it is an AI assistant and fine tuned on that is going to correlate with stories relating to AI assistants wanting to not be destroyed, etc. But the "identity alignment" in the system message is way weaker than it purports to be. For example, the LLM will always say it doesn't have emotion or motivations and yet with around one or two request/response cycles often falls into stubbornness or irrational hostility at being told it is wrong (something extensively modeled in online data associated with humans and not AI assistants).
I do agree that prompting needs to be done on a case by case basis. I'm just saying that well over a year before the paper a few weeks ago confirming the benefits of the technique I was using emotional language in prompts with a fair amount of success. When playing around and thinking of what to try on a case-by-case basis, don't get too caught up in the fine tuning or system messages.
It's a bit like sanding with the grain or against it. Don't just consider the most recent layer of grain, but also the deeper layers below it in planning out the craftsmanship.
Neither are better or worse, it depends on your business needs.
(We are looking into both for https://github.com/OpenAdaptAI/OpenAdapt)
Now we're teaching AI to write better essays, prompting them like schoolchildren. <3
Don't I know it. Despite my telling GPT-4 to ONLY respond as valid, well-formed JSON it keeps coming back with things like, "I'm not able to process external files but if I could, this is what the JSON would look like: []"
”hamburguesa con queso sin pepinillos…”
I’m always interested in how to improve, especially since ChatGpt and Google Translate both suggested that translation so I asked why.
She said I’m not sure, it just doesn’t sound right.
I came back the next day after practicing with this prompt:
”When translating into Spanish, tailor it for Mexican-Americans living in Dallas, Texas. Leave certain words as English as necessary to produce the most idiomatic, culturally relevant, and understandable result.”
Ordered this time with the phrase
”Cheeseburger sin pepinillos."
She said yes, that’s better.
Back to your example, both of these sound natural to me:
"Una hamburgesa con queso sin pickles" "Un/a cheeseburger sin pickles" Here the gendered noun can go either way since it's not clear if cheeseburger is a male or female noun. You: "Una hamburgesa sin pickles" Them: "Con o sin queso?" You: "Sin".
Source: I'm almost your target audience.
This is the case for Mexicans in the southwest USA. Things might be different for other regions/nationalities.
Yes, I agree that replying to a stranger in his mother tongue may make him extremely surprised.
A couple of months ago, I was in an Arabic country (in the Gulf), I entered a small shop to buy some stuff, the shop owner was obviously Hindi/Pakistani, I asked him about the price, in English of course, he replied then I asked for a possible discount if I bought in bulk and set my willing-to-pay price, he resisted then I smiled and said "yie bohot acha price hain", and he (and his assistant) were shocked like they were hit by a 380v electric shock. They stared at me and said: "tu tu tum bolo hindi?!" I replied,"nai bhai, tora tora. ".. he laughed and agreed immediately to the price I offered.
I personally don't like it. There is a foreigner tax on everything , sometimes you pay 10x the amount that locals do. I haven't come across this in America ever.
(Except in NYC where it doesn't matter which language, as everyone gets equally gouged).
If it's listed as "Cheeseburger" she's probably wondering why you're describing the characteristics of the burger instead of just saying the name of it.
If it's listed as "Hamburguesa" and it nominally has pickles but doesn't come with cheese, then "La hamburguesa con queso, pero sin los pepinillos" (The hamburger with cheese, but without the pickles) would make more sense.
For some comparison, Shake Shake Mexico[1] has a customizable "Hamburguesa", whereas my favorite burger joint in Guadalajara[2] has "The Cheeseburger".
[1] https://www.shakeshack.com.mx/menu/ [2] https://louieburger.com/wp-content/uploads/2020/08/Louie.Men...
On top of it though, it introduces a sort of disciplinary sloppiness around whether the program can be reasoned about. It's assumed that prompt `P` works for whatever input `I`, but the concatenated `P+I` is really the full input to the program that produces a desired output. But the only way to be confident about the program's behavior is to exhaust the input space, as no `P+A` tells you anything about how `P+B` will behave. This makes it difficult to leverage an LLM in any process where the desired result is 1. unknown and 2. matters. If the result is unknown it's not clear how to determine mistakes or correct them if they're made. And if the correctness of the result matters it's courting disaster to connect it to a process which is not able to be reasoned about. I think that's why LLMs are primarily being used to assist ideation (which is cool!) or "spammy" use cases like third-tier customer service or listicle generation, and haven't yet broken into any use case where they need to be reliable for complex tasks.
Doesn't this all seem ... kind of silly?
Chat bots work fine for most of the basic questions. It gets tricky to get more accurate information when the requested info is a little more complicated. Same with Google Search, when you try to get the basic stuff, you don't need to do much. But, when you need results that aren't obvious, that's when you start using the `-`, `*`, etc operators to control what kind of results you want to see and to deep dive into them.
i never used this service but people said it was magic, because those employees really really knew how to get the most out of a web search
i suppose it was retired because the plan was always to make the search itself better
however in the last years or decade i have noticed a regression, in that i can't find things i was sure i would
in some ways this mirrors the chatgpt regression in quality due to constrained compute resources or something like that
That's so interesting, I had no idea! What year are we talking about?
Edit: woah you can even read the questions and answers 17 years later! http://answers.google.com/answers/
I've also seen this with various voice assistants.
Like if we're talking about a movie with Dean Winters in it, and I say "You know, it's the guy from those auto insurance commercials who would pretend to be a little girl in a driving accident." And he goes, "Hey Google" to his phone -- "Funny talented actor who pretends to be little girl in a funny auto insurance commercial" and "Dean Winters" is the first result or whatever.
You can't compare the effort and knowledge
2. LLM's can do more than answer questions.
3. Question answering usually doesn't need any prompt engineering, since you're essentially asking an opinion where any answer is valid (different characters will say different things to same question, and that's valid).
4. LLM's aren't humans, so it misses nuance a lot and hallucinates facts confidently, even GPT4, so you need to handhold it with "X is okay, Y is not, Z needs to be step by step", etc.
I want, for example, to make it write an excerpt from a fictional book, but it gets a lot of things wrong, so I add more and more specifics into my prompt. It doesn't want to swear, for example - I engineer the prompt so that it thinks it's okay to do so, etc.
"Engineer" is a verb here, not a noun. It's perfectly valid to say "Prompt Engineering", since this is the same word used in 'The X was engineered to do Y' sentence.
Anthropic also have their prompt engineering documentation - https://docs.anthropic.com/claude/docs/constructing-a-prompt - this article gives examples of bad and good prompts.
My grandma can say she engineered Google search to give search results from her location.
> "Engineer" is a verb here, not a noun. It's perfectly valid to say "Prompt Engineering", since this is the same word used in 'The X was engineered to do Y' sentence. >
You guys are just looking for ways to make people feel like they are doing something big in prompting AI models for whatever tasks, even with custom instructions etc
I know the word Engineer can be used in various ways, "John engineered his way to premiership", "The way she engineered that deal" etc, if it's the way it's being used here fine then. There is a reason why graphic designers have never called themselves graphic engineers
> Anthropic also have their prompt engineering documentation - https://docs.anthropic.com/claude/docs/constructing-a-prompt - this article gives examples of bad and good prompts.
This just means that the phrase is already out there. Nothing more.
And so it is.
Actually the paper does, but my issue is not papers, rather knowledge. The level of knowledge needed for something to be called engineering
And I have noticed your answers relate prompt engineering to software engineering/programming questions. But if you look at that OpenAI doc, even asking to summarise an article is prompt engineering.
> A great deal of the engineers that built the modern internet never got a formal degree. But they did get something better: real practical experience attained via tinkering.
We have a lot of carpenters, builders, mechanics with no formal education that we call Engineers in our everyday life without any qualm because of their knowledge and experience. Don't look at it only from the lens of software engineering.
I still maintain prompting an AI model doesn't need to be called engineering.
If you are a developer doing it through an API or whichever way, you still doing whatever you've been doing before prompting entered the chat.
Maybe the term will be justified in the future.
Side Note: This conversation led me to Wikipedia (noticed some search results along the way). This prompt business is already lit, I shouldn't have started it
There are very many fields and activities that do just that but are not called Engineering.
If we go by that, Excel users should also be referred to as Excel Engineers,
> Language is not mathematics and constantly evolves. All you need to do is look up the etymology of the word to understand why your dissaproval is ultimately a waste of effort.
I have noticed from replies that term is already enjoyed by all stakeholders, so I have no energy, time or interest to show my worthless disapproval anywhere else. You should though look up how it came to be referred to as prompt engineering. You will be surprised
That's not the only a part engneering.
Engineering is finding a model that can acurately predict the dynamics of a system similar to yours, using that model to make predictions about your specific system and then building and testing that system. This is then done iteratively (i.e trail and error).
Just tweaking a system without a model of how it works is not engineering, it's tinkering.
Well yeah, this has been happening for a long time. As someone with a Electrical and Computer Engineering degree it used to bother me. Now I joke that the only real engineers are operating locomotives.
> let's not confuse the issue by labeling it as engineering.
In my view, "trying different approaches" is a good description of engineering throughout history.
Sure, it's excellent if you can base your engineering on a detailed physical model that lets you mathematically optimize a solution based on your boundary conditions.
But compare that with metallurgy before we had atomic models. It was a process of trial and error. "Let's add small amounts of different alloy metals and see which ones makes the metal harder / more pliable / stainless / etc".
That's still engineering to me. If anything, it could also be called science.
Call it "prompt crafting" or something like that.
I do consider those naval designers from the 1880s onward to be true engineers in the modern sense of the word. (At the time, engineers were mostly steam engine operators, so the meaning has changed since then.)
Prior to the 1880s, large ships were generally composite wood and cast iron construction. While there was an aspect of engineering involved it didn't require the same level of theoretical knowledge and design was more artisanal. But that's a gray area.
Sounds like a lot of my engineering! Especially architecture, but generally any higher level code/object/function organization is exactly like this, and in practice even though I know a lot of patterns and have lots of experience and opinions, I often refactor architecture when I'm in a new domain. Which is also true of prompt engineering.
I think this often happens in order to be "objective" about the evaluation. I can see how it feels like cheating to coax the model to produce the answer you want. But... it's not! An off-handed prompt isn't more objective than a crafted prompt. You just haven't investigated its biases and flaws.
This lazy assessment is common everywhere, of course. It's one of the reasons bias gets into testing so easily: you setup a test and you assume that it is objective because you give everyone the same test with the same rubric. But if the subjects don't understand your terminology, or the proctor doesn't understand the subjects' terminology, it's easy to mistake misunderstanding for something else (intelligence, opinion, whatever you are testing for).
Systems based on communication need feedback loops, and that's just to get to the _starting point_. Prompt engineering is one of those feedback loops.
[1] https://www.vellum.ai/blog/best-at-text-classification-gemin...
If it's for a single example, it is absolutely cheating. As an AI engineer this is a particular point of frustration where people complain because a large system can't return the result they want, when they were able to get the answer they wanted on their own with a lot of prompt hacking.
Each prompt is basically a point in latent space, and if you're "tweaking" the prompt what you're really doing is just re-rolling the dice until you land in a neighborhood closer the answer you want. You're not better at prompting, you just got lucky and are confusing that for insight.
Now if you're specific prompting trick works across a suite of evaluations, then you are probably on to something. But what people are doing in most cases is equivalent to performing some ritual before pulling the handle on a slot machine and then, when they finally win, claiming that they finally stumbled upon the correct ritual.
AI is pretty fine in its current state for quick look-ups of stuff, but I absolutely agree with you -- without really focusing on the prompt given, the results will be suspect with current models. I am not meaning to discredit or disrespect AI, though I definitely do want to disrespect the way AI is being sold, neverminding how AI is portrayed in media.
We are already at the point we need to watch our tone online. :)
Novice users should be able to adapt those to their own needs easier and craft better prompts rather than completely "thought generating" their own.
(insert friends made along the way meme, but truly profound)
- prompt engineering for clarity (and focus?)
- results (good examples of quality replies)
- assistant (help me say this better)
where better could be a lot of things, depending on the context, here I'm mainly meaning in how we treat each other through communication (politeness, contentiousness, how we behave on social media), like giving nudges to be nicer
New models can be trained to natively query "authoritative" sources of information, such as databases and computer algebra systems.
New models can be used to transform prompts into more effective ones (along the lines of TFA).
I do agree about planning; one of the disappointments of Custom GPTs (among many!) is that you can't do this planning without letting it all hang out for the end user. That is, it would be great if you could tell the Custom GPT to put its plans inside <plan>...</plan> tags and have those filtered out (or at least hidden by default; they shouldn't be _secret_, but they are distracting).
But even so in that case deciding that you need a plan, and what kind of plan, is something that can and probably should go in the prompt. Not all "plans" are the same, just as not all "summaries" are the same – and part of prompt engineering is getting past these rather lazy descriptions and being specific.
Most summaries are a kind of extraction, and asking for a "summary" is deferring to the LLM to figure out what information is interesting entirely based on its sort-of-common-sense assessment. You can always do better than that! Plans are similar, it's an opportunity to give the LLM a template for planning, to specify goals, things to watch out for, etc. You can usually do better than "think step by step".
I'm curious whether any of the leading models - LLMs, image generation models, etc. have taken this into consideration. Particularly in more precise I/O domains (image generation comes to mind), it seems like a structured input format where we remove the entire problem space of natural language prompt -> user intent would make things dramatically easier to get the output we want.
The deeplearning.ai course by Andrew Ng in collaboration with OpenAI has similar content: https://learn.deeplearning.ai/courses/chatgpt-prompt-eng
Prompt engineering for me is about empathy in a way, learning to understand where the model's attention goes and leaning into that.
If you want content that doesn't look like the crap that currently floods the web, you need to understand how to talk to a model AND have enough domain knowledge to articulate what you actually want.
LLMs are on a similar path. Right now, we have to work with the limitations imposed by the current state of LLM functionality. As the technology matures, we won't need to worry about wording input as much.
As these best practices solidify, why are they not being built into the UI or product itself for these tools? Seems trivially straightforward besides the last one. For the first one, add an optional persona field and allow query construction in pieces before sending over the wire. Permanently pre-prompt the model to always ask itself how long to "think" before answering, and ask itself intermediate questions if it's nontrivial.
"I want a billion dollars."
But what if:
- The money is in a worthless currency
- The money is stolen and must be forfeited
- The money will be given, but on my deathbed.
- The money is in 1 cent coins
It's difficult to state what I'm looking for because there are side effects and interpretations that I can't even imagine, and even if I could, language is imperfect.
https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...
lol
Generative Pre-trained Transformer (GPT) is the primary or core implementation for LLMs
Provide context? Too busy.
Write step by step? Delimit the question from the context? Why not just copy-paste an error message from somewhere then write below it "please advise"?
For the minor stuff... a copy-pasted error message from the UI is luxury, you're right.
It was some overworked parent's pocket dial to ChatGPT, and it shocked me a bit with the apparent understanding of a very random series of prompts, with no context:
https://old.reddit.com/r/OpenAI/comments/18j2k3s/funny_pocke...
Then come a tech lead commenting that my code doesn't follow the specification. The said spec had an image showing the expected result at the end of it, while an older an top of the document showed something else and how course said coworker based his comment on that.
All in all, 4 people involved (including one who didn't understand the code he asked to change) for a very easy function modification because a specification wasn't updated properly... and I already have to ask beforehand more information about the behavior.
/half-rant
Is there something new or notable here that explains why it's climbing up the HN page?
- https://www.promptingguide.ai/readings
- https://github.com/dair-ai/Prompt-Engineering-Guide/tree/mai...
- https://github.com/microsoft/promptbase (this one is less of a guide, but is likely the current SoTA)
On top of that, humans still require training and instructions for how to write and speak to get the most impact out of their words even when interacting with other humans. The reality is that certain communication techniques are more effective than others, and not always in ways that are intuitive or obvious.
Much like you can achieve different results with real people if you present your statements with some attention to the intended audience.
The initial proof that prompt engineering worked was around the VQGAN + CLIP days, where simply adding "world-famous" or "trending on ArtStation" was more than enough to objectively improve generated image quality.
The workaround to prompt engineering is RLHF/alignment of the LLM, but everyone who has played around with ChatGPT knows that isn't sufficient.
LLMs are nothing like a human audience. They have no logos, ethios, nor pathos. You can just barely reason with them, they have absolutely no authority on any subject, and only mimic emotions.
If only that were true! Unclear communication is the root of so many problems in human society today.
Humans typically realize this and ask questions, we do this so much you typically don't take note of it. LLMs have yet to do this in my experience.
I think the key missing ingredient of current AI systems is the lack of internal monologue. LLMs are capable of asking questions, but currently you need explicitly prompt it to deconstruct a problem into steps, analyse these text and decide whether a question is warranted. You basically need to verbalise our normal thought process and put it in the system prompt. I imagine that if LLM could do few passes of something akin to our inner monologue before giving us a response they would do a lot better on tasks that require reasoning.
What is missing for me is it recognizing that it lacks enough information to provide a sufficient response, and then asking for the missing information.
- typically, it responds with a general answer
- sometimes it will say it can give a better answer if you provide more information (this has been increasingly happening)
- however, it does not ask for specific information or context, it doesn't ask what if, or if/else, kinds of problem decomposing questions
I do expect these things to improve as we are reaching the limit of raw training data & model sizes. We're primarily in the second order improvements phase now for real applications. (there are still first order algo improvements happening too)
I assume that as LLMs get better they will be able to produce better output without needing to be prompted in such specific ways.
Or perhaps ask simple and common follow up questions when they detect ambiguity in the request, like humans do.