Prompt Engineering Guide: Guides, papers, and resources for prompt engineering
github.com
github.com
From Prompt Alchemy to Prompt Engineering: An Introduction to Analytic Augmentation:
https://github.com/williamcotton/empirical-philosophy/blob/m...
A few more edits and it's ready for me to submit to HN and then get literally no further attention!
Here's another approach of second-order analytic augmentation, PAL: https://reasonwithpal.com
Third-order, Toolformer: https://arxiv.org/abs/2302.04761
Third-order, Bing: https://www.williamcotton.com/articles/bing-third-order
Third-order, LangChain Agents: https://langchain.readthedocs.io/en/latest/modules/agents/ge...
The difference isn't in what is going on but rather with framing the approach within the analytic-synthetic distinction developed by Kant and the analytic philosophers who were influenced by his work. There's a dash of functional programming thrown in for good measure!
If anything I filled a personal need to find a way to think about all of these different approaches in a more formal manner.
I have scribbled on a print-out of the article on my desk:
Nth Order
- Existing Examples [x] (added just now)
- Overview []
- data->thunk->pthunk []This is currently how a number of existing tools work, both in published literature and in the wild with tools like BingChat.
I have personally found this analytic-synthetic distinction to be useful. I'm also a huge fan of Immanuel Kant and I really love the idea that I can loosely prove that Frege was wrong about math problems being analytic propositions!
They were recently given $300M by Google, so certainly you have a promising idea there
I'm definitely not calculating the path of a projectile in a similar manner when I catch a ball.
I'm definitely not "computing" a sentence when I read it in the same way that I compute the multiplication of those two three digit numbers!
You're not calculating it in a traditional sense, but there's definitely some systems of partial differential equations being solved in real-time.
I was increasingly frustrated with all the NLPness and operator deprication in google which has been accelerating since at least the 2010s.
But with Kagi it reallys makes me feel like I am back at the wheel. I think for product search it still has some way to go but for technical queries it is just on a whole other level of SNR and actually respects my query keywords.
Isn't this just an intrinsic problem with the ambiguity of language?
Reminds me of this: https://i.imgur.com/PqHUASF.jpeg
edit: especially the 1st and last panels
Prompt: "Calculate the average price of Milk"
This is far too vague to be useful.
Prompt: "Calculate the average Price of Milk between the years 2020 and 2022."
A little better but still vague
Prompt: "Calculate the average Price in US Dollars of 1 Gallon of Whole Milk in the US between the years 2020 and 2022."
Is pretty good.
For more complex tasks you obviously need much more complicated prompts and may even include one-shot learning examples to get the desired output.
This argument is weak. Undefined behavior does exist and „high level programming language „is a moving target
Other than complicated requests like “if object is of type A, include fields ABC but not D. If object is of type B, include only D but not other fields”, it gets this right 99% of the time.
It also works for CSV, but it’s trickier. It seems like it “knows” how JSON works to a much better extent.
And as for parsing JSON? I’ve not truly pushed it to its limits yet, but so far it’s had no issues understanding any of it.
It’s mind-boggling. Yes, it’s inefficient, but it can basically parse, generate and process valid JSON with just a brief set of instructions for what you want to do. For exploring ad-hoc data structures or quick mocking of API backends, this is great.
2. the action of working _artfully_ to bring something about. "if not for his shrewd engineering, the election would have been lost"
(https://www.google.com/search?q=define%3AEngineering)
Merriam-Webster has:
3 : calculated manipulation or direction (as of behavior)
giving the example of “social engineering”
(https://www.merriam-webster.com/dictionary/engineering)
Random House has:
3. skillful or artful contrivance; maneuvering
(https://www.collinsdictionary.com/dictionary/english/enginee...)
Webster's has:
The act of maneuvering or managing.
https://github.com/dair-ai/Prompt-Engineering-Guide/blob/mai... (this is claimed for LLMs, not proven)
It only seems like a trick until enough papers get written about these kinds of findings.
This includes asking questions, and trying to direct someone to effectively complete a task.
Prompt engineering is just communications skills as applied to AI instead of meat-minds.
E.g when talking with my accountant, she usually ask me a bunch of clarification questions and numbers, instead of just making a best effort, but confident sounding response with whatever initial context i gave.
I myself don’t have to have the depth of knowledge to direct someone to complete a task step by step to get the right results.
Perhaps a big step forward for chat AI interfaces is to make the AI capable of knowing when it needs to ask follow-up questions and then having it do so. Essentially, it helps you along with your prompt engineering.
Perhaps this is what the prompt engineering tools are doing anyway.
What we need to do is integrate the prompt engineering tools with the chatbot itself, so it can both help extract a good prompt from users and then answer that prompt in the same process.
I think this is where we'll move towards relatively soon; it seems obvious when you say it.
But they don't.
If you did have to send google-esque one shot queries to humans you'd probably settle on short hand like "polish shiny" or you'd opt to use "poland" specifically to avoid it. For most known-ambiguous terms you'd do this, the same way we say out loud "M as in Mancy" because we know the sound of the letter can be ambiguous sometimes. We have lots of these in English that you use all of the time without knowing it. In a smaller audience the local language gets even more specific, you probably disambiguate people with your friends like "programmer Bob, not carpenter Bob". It's not at all crazy that we'd develop another local language for communicating with computers even if that's not a traditional programming language
If I were talking to a python programmer I could assume they knew what a for loop was, so could phrase a question requiring the context of one differently than I would for the layperson. Just like if I were designing inputs to one language model I could assume it's capabilities that are different from another.
I don't think it's relevant to the performance of language models for that same reason though, we already have to design out queries with the thing being queried in mind so I don't see why we wouldn't for llms.
Even the becoming-cliche "humans don't know what they want" argument is its own counterargument. You're right, humans do not know how to precisely ask for what they want, and yet manage today to navigate around this all the time without learning a new way to communicate.
But, that does not close off the possibility of learning what ways we can poke and prod it to affect the results.
It would be interesting to run an experiment to test whether some people are consistently better at generating desirable results from an AI model. My money would be that they can, even at this early stage of "prompt engineering" as a discipline, let alone 5 or ten years from now.
You may also be saying "don't call it engineering, it feels more like black magic," a position I would be sympathetic to. But, I think a lot of realms of engineering deal with uncontrollable elements we don't understand and have to just deal with, with increasing levels of control as our disciplines become more sophisticated.
But what is your suggestion? I feel it would be recieved easily if it was LLM Prompt Cheat sheet or something, though I wouldn't have seen it on HN.
I don't see why you find that controversial.
Translation is a skill, and translators remain employed today despite ml improvements in automated translation over the years.
Humans do talk in different variations of language for different tasks. Eg. Legalese and code switching.
1. Currently, Prompt Engineering works, and that alone is a reasonable reason for people to explore it. The concept of doing nothing about things that work today because there might be a day when they become obsolete makes no sense at all.
2. Prompt Engineering is important and meaningful. Most people are thoroughly incompetent at giving instructions. And it's a core skill of projects. An incompetent person will fail, no matter how competent their subordinates are.
3. PE is something that will be needed even as AI gets better (until we have omnipotent superintelligence). In fact, even giving instructions to humans requires PE. It's a limitation of human language.
4. I think it's also wrong that people's motivation for exploring PE now is an attempt to make room for humans. Have you ever been to an AI drawing community?
5. It's not PE that should exist for human intermediaries, it's the skills to handle AI, including PE. If humans no longer need the skills to deal with AI, it will be AGI by definition. If what you're saying is that when AGI comes along, all humans will be unnecessary, that might be true, but then what you're saying is meaningless.
However, I think this kind of 'engineered prompt' sharing is about as useless as sharing the engineering calculations for building a bridge over a particular river : none of those calculations are generalizable and all must be redone for each future placement.
Chances are you'll need to tweak any purchased prompt to fit your own use cases, and if you are capable of tweaking purchased prompts, why can't you just chat with the bot for an hour or two and build a working prompt for yourself?
I think the value of "prompt engineering" as a skill is all about adaptability and testing and verification processes: what is the process of going from idea to functioning prompt for a particular use. Silver bullet prompt fragments are cool, but not really products - the only people buying these are suckers who don't understand how to use chatGPT. Because chatGPT will happily help you make these prompts for free.
Honestly, my favorite moment has nothing to do with the thesis of the video: it is when she says "zero-shot chain of thought" and stops talking for a moment to comment about how that rhymes, as it definitely only rhymes because of her accent; but I'm really into linguistics and, in particular, phonology, so, for me, that kind of moment is magical.
But like, the point is: if you watch that video and replace the idea of you talking to ChatGPT with talking to one of your coworkers, this is a totally legitimate video someone would make. You can't quite just use a book of brainstorming ideas, though, with ChatGPT, as it has some fundamental limitations to how it understands certain kinds of concepts or language, and so seeing "oh that path worked well" is practically useful.
I don't doubt that there is value is sharing useful workflows to learn how to write prompts or useful things to add to prompts, but I've seen a lot of people selling a 'list of tuned prompts' as if each was a plug-and-play deployable software stack. I think anyone who tried to actually productize these prompts would find the need to tweak the prompts so much they might as well have started from scratch.
In short, I think there is value in teaching people how to write prompts to custom-fit their needs, but the idea of general purpose couple-paragraph superprompts that mere knowledge of is worth tens of dollars seems fallacious to me.
Disclosure: I've never bought any of these, maybe they are that good. I'm very skeptical though.
https://www.theverge.com/2023/2/2/23582772/chatgpt-ai-get-ri...
But working of the fables of the genie, we have and all powerful wish machine where the user mistakes they are telling the genie with intention of their thoughts, whereas the genie is looking at all possible decodings of the words they are saying without knowing the users thoughts. If you do not verbalize exactly what you want (and have done some thinking about the ramifications of what you want) you might end up squished under a pile of money, or some other totally ridiculous situation.
Unless you share your entire life with the AI/genie you will always have to give more detail than you think. Hell, in relationships where people live together for years there is always the possibility for terrible misunderstandings when communication breaks down.
Scale.ai (no affiliation, not a customer) has a product called Spellbook
I'm wondering if we'll reach a point where it is impossible or nearly impossible, and what value of 'understand' we assign to it. For example deeply understanding x64 processor architecture. Does any one person understand it all, very unlikely at this point. The investment the average person would have to perform to accomplish even part of that would lead most of them to seek other more fruitful endeavors. Biology would be a good example of this. Nothing about biology is magical, and yet even simple systems contain such an information density that they are at the impossible to fully understand point, and that's before we start layering them into systems that have emergent behavior.
In this case though, the increasingly desperate engineering of the prompt was each time intentionally side-stepped by the devil, though much like GPT, he couldn't help being like that, it was just his nature.
We need a name for this, something like genie effect
Investigating why the prompt isn't working:
"Debazzling"?
As for some philosophical inspiration, the analytic philosophers like Kant, Frege, Russel and early Wittgenstein have methods for breaking down natural language into (possibly) useful elements!
Like, everyone speaks of "context" in terms of these prompts... how similar is that to Frege's context principle?
https://en.wikipedia.org/wiki/Context_principle
Some other Wikipedia links:
https://en.wikipedia.org/wiki/Linguistic_turn
https://en.wikipedia.org/wiki/Philosophy_of_language#Meaning
It feels like tutorials on tailoring the emperor's new clothes.
At the same time I feel it's only a matter of time until my job interviews involve questions like "can you provide a few example of where you've used prompt engineering to solve a problem? Describe you prompt engineering process?" and I'm just not sure I can fake it that hard.
One for context, one for rules, one for output format, one for input data, etc.
Also, you can as for a structured output (gpt spits out JSON fine) but you need to be super explicit about each field possible values and the conditions they appear in.
Templating prompts with jinja goes a long way for testing.
And you will need a lot of testing, if for nothing else than remove variability in answers.
It's funny because people keep saying gpt will removing all values from having writer skills, but being a good writer helps a lot with crafting good prompts.
So I've been writing the translation examples (few-shot examples [not my favorite term]) as TypeScript, transpiling to JS and compiling with other examples and a prelude, and using the same method to build the translation examples as to build future prompts. It saves a lot of silly mistakes!
JavaScript has less tokens than TypeScript, so it seems more economical to convert to JS beforehand instead of passing TS to the LLM! I wouldn't be surprised if TS resulted in better solutions, though... add it to the endless list of things to test out...
This is pretty similar to the approach used in the Toolformer paper, other than sample-and-vote, which I believe is novel.
Toolformer: Language Models Can Teach Themselves to Use Tools
https://arxiv.org/abs/2302.04761
The authors show a valuable method of using prompts to generate training data to be used for fine-tuning.
Check out Appendix A.2 in the paper for the example prompts.
As an aside - does anyone have good tools or methods for testing and evaluating prompt quality over time? Like performance monitoring in the web space, but for prompt quality. The techniques to use LLMs as evaluation tools of themselves always has seemed flakey when I've tried it, I'd like to use a more grounded baseline.
For example, if you have a prompt that says "What is the weather today in {city}?", you can run it against a list of cities and expected outputs (using a lookup to some known truthful API). That way, when you make changes to the prompt, you can compare performance to a baseline.
Some more contect - as a way to support MakerDojo[1], we are building TryPromptly[2] - a tool to do better prompt management. In that tool, we are building the ability to create a test suite, run the test suite and compare the results. At least knowing for which test case the results varied and reviewing them would go a long way.
Here is the format we are thinking: https://docs.google.com/spreadsheets/d/1kLBIb7W0jrY-IkNPqJsN...
In addition, we are about to launch after test suite is to have live A/B tests in the production based on user feedback. Users can upvote or downvote their satisfaction with the results and that will inform you which version of the prompt yielded better results.
If you have other ideas on how to test them better, it would be super helpful to us.
[1] MakerDojo - https://makerdojo.io [2] TryPromptly - https://trypromptly.com
Human like AI isn't going to be magic because humans are not magic. You are still going to have to comprehensively communicate with your AI so it understands your reference frame.
If you walked up to the average programmer on the side of the street and threw a programming fragment at them, especially if its a difficult problem, they will have a whole load of follow up questions as to place the issue you want to solve. Coming at a human or AI with expectations and limitations first will almost always lead to a faster and better solution.
If there's anyone here who's become good at perfecting llm prompts & available for a freelance contract -> please contact me (my email is in my profile)
It's easy-ish to get gpt to generate "good enough" text, obviously. Like any tool, what's interesting and complicated are the use cases when it doesn't follow instructions well / gets confused, or you're looking for a more nuanced output.
Composing the right prompts here will allow you to obtain actually what you want versus, let's say, a castle as drawn by a toddler (okay maybe that's what you want, but that's not the point)
For example, prompts that do not contain enough facts about a desired question will result in "hallucinations", whereas prompts that have additional context programmatically added to the question will result in more reliably factual completions. I find it useful to think about this grounded in terms of the analytic-synthetic distinction from analytic philosophy.
The above guide is an excellent resource! You should read it!
https://medium.com/data-science-at-microsoft/building-gpt-3-...
In short, Prompt Engineering is just one of the pieces in building a GPT-3/LLM-based solution. In fact, I'd say a whole new set of Software Engineering best practices is necessary and will gradually emerge. I gave one such approach that has been useful to my related projects.
But in this universe, it took a different turn. Words are a funny thing. Ironic.
42?
Why is it that we create something with which we can't communicate but that we expect to solve all our problems.