Promptbase: All things prompt engineering
github.com
github.com
I'm also curious as to why existing LLMs can't be fine tuned to handle this, if prompt engineering is really a major concern.
Usually, if I don't get the answer I'm looking for from ChatGPT, I tell it it's not the answer I'm looking for, what the original answer was missing, and I usually get a better answer the second time around.
If it goes beyond that I sometimes resort to just cussing at it, and that usually does the trick.
I've been curious if DALL-E has been truly been mixed in w/ ChatGPT as a single model using a Mixture of Experts (MoE)- type learning gate to train them all together.
This is an expected behavior, as by default even with RLHF ChatGPT will output statistically "average" content. It's also the reason why Chain of Thoughts prompting works very effectively.
I have a (somewhat out of date) notebook demonstrating this, plus a function calling trick which allows you to get the improved result in a single API call: https://github.com/minimaxir/simpleaichat/blob/main/examples...
They can.
It's just important enough for performance to be worth doing this manually.
> I'm also curious as to why existing LLMs can't be fine tuned to handle this, if prompt engineering is really a major concern.
It depends what you want them to do. They're general purpose things and there isn't a one size fits all solution here.
I guess in the 2nd case it would be useful to ask it to output the reworded prompt too.
If your app or feature is powered by an LLM, why pay for the state-of-the-art LLM when you can get comparable performance out of a cheaper one with a little bit of prompt engineering?
Over time, as the models themselves improve, prompt engineering may become less important. I already think it's largely unimportant for day-to-day one off things.
- https://chat.openai.com/g/g-LQHhJCXhW-autoexpert-chat
Phind has a whole Discord channel and they're pretty focused on building a great tool aimed at programmers. It's approaching pinned-tab status for me.
Disclosure: I work for Microsoft; this advice is my own.
You are an expert JavaScript programmer. Write a JavaScript function based on the user input.
You must obey ALL the following rules:
- Only respond with the JavaScript function.
- Never put in-line comments or docstrings in your code.
And then editing iteratively based on the output to whatever your desired use case is. If you want to test ChatGPT system prompts directly in a UI, you can do that in the OpenAI Playground: https://platform.openai.com/playground?mode=chatI've found that GPT-4 works pretty well if you just talk to it like a person.
[0] Prompt Engineering Guide (https://www.promptingguide.ai/)
I am not really sure how to answer this question, seems like a bunch of skills.
Even from an etymology standpoint engineer comes from the Latin ingenium meaning clever.
that is a portion of what I have been referring to.
The real problem is that we insisted on giving a discipline that is far more art than science the title "engineering".
I interpreted the original comment to mean "I won't consider ChatGPT an artificial intelligence until we don't need to prompt engineer." If that was the intended meaning, I just wanted to highlight that we do "prompt engineer" humans while also considering them "intelligent."
I don't know what "prompt engineers" think they're "engineering." There's nothing of the sort remotely happening here. This is just random uninformed actions being tested against a weak fitness function. The results are effectively meaningless in any broader context.
https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...
It's almost 2024 and I still can't ask ChatGPT for "the definition of {apple|orange}" where {apple|orange} is the mathematical embedding average of those two words. I can sure do this in Stable Diffusion though!
You can do prompt term weighting with offline LLMs (e.g. compel: https://github.com/damian0815/compel ) or averaging the embed tokens yourself and passing the embedding matrix to the model, but unlike with image generation where the results of prompt weighting are more obvious, there isn't as much of a need for it for LLMs.
I shouldn't have to "average the embed tokens myself".
I know you're a big deal in the industry, but you're SO WRONG about the idea that "there isn't as much of a need for it for LLMs". I hate that the NLP community has such huge blindspots for this kind of stuff. Anything that gives further levers of control has massive improvements to the capabilities of an LLM.
My github gist proves that all of these techniques found in automatic1111 work with NLP models, yet no one implements it. I think it's because of the potential for breaking alignment techniques with them.
And yes, I will fight and die on this hill even if Yann Lecun and Christopher Manning tell me I'm wrong. I know I'm right.
No it doesn't, your gist just shows that it's possible to implement, which I'm not disputing. And even then, your example with GPT-2 has the comment "GPT2 is very tempermental with this technique" and requires a temperature of 20 to behave which is an accommodation that can never be used in a real application.
If you can create test cases and demos where prompt weighting in a LLM results in a distinct observable improvement as it does with diffusion models, that would be a different story. I'd love to be proven wrong but your gist doesn't do it.