GPT-Prompt-Engineer
github.com
github.com
Why is this so popular, then (more popular than promptfoo, which I think is a much better tool in the same vein)? AI devs seem enamored with the idea of LLMs evaluating LLMs —everything is ‘auto-‘ this and that. They’re in for a rude awakening. The truth is, there are no shortcuts to evaluating performance in real world applications.
Here is what HN was talking about, nearly three months ago -the exact same type of ‘auto-prompt-gen’ tool: https://news.ycombinator.com/item?id=35660751
I was reminded of the same thing. What a lot of it boils down to is that LLMs have no innate ability to self-reflect. They can pretend to do it, but no more effectively than an untrained human would.
Which is exactly as much as Generative AI should be trusted.
If you're prepared to accept that GPT-4 can answer questions just as well as humans can, why do you even need to do prompt engineering?
* What's the best way to get to Radio Shack from here?
is not the same as
* What's the easiest way to get to Radio Shack from memory when riding a bicycle from here?
You may be under the impression that annual U.S. deaths from medical errors being in the hundreds of thousands miscommunicates but that is truly your opinion. You are merely jumping to conclusions at places another person might not.
And going on to rely on the LLM to validate your perspective is a lossy process. It may not lose your perspective but it loses someone else's and you don't even seem to notice or care.
The post you replied to was saying that the deaths were caused by miscommunication, but you interpreted it to mean that stating the number of such deaths is somehow a miscommunication itself!
Yeah I agree there. Unless you can check against the output, it’s not really telling much.
Could you post the link please
The problem is that it will simultaneously say that "cow eggs are bigger than chicken eggs", with the same confidence (and in a way that correlates well with human evaluators).
https://www.reddit.com/r/Funnymemes/comments/10ohd2n/chatgpt...
So when you get an evaluation you are playing the russian roulette - you may get a decent result, or you may get cow eggs.
The point is that the tool fails, and it is known to fail, so much that we even have a name for the times when it fails - hallucinations. I have been calling them cow eggs because that's a nice mental image and I didn't want to have to remember for the proper English term. I will continue calling them cow eggs.
A positive correlation just means better than chance. "Strongly" is vague, and might not be much better than chance.
No, adverbs like “strongly” modify adjectives (or verbs, but that’s not relevant here) not nouns; “strongly” is an intensifier that modifies “positive”, its not a separate adjective that modifies the noun “correlation”.
You can't use the thing you're testing to evaluate its own performance. This applies to rulers, speedometers, and AI. It's the difference between a "subjective" and "objective" metrics. If you want an objective metric, you need to have it based on something external, based on reality, objective. Otherwise, you have metrics and ideas that have to held themselves up.
Source: My day job is test and measurement. These concepts go back centuries. You never trust your measurement system, you verify it against a standard.
Grifters. I won't say that the person working on this is a grifter. Instead, it's so popular right now because of grifters. The same type of NFT grifters and crypto grifters who are mostly silent now. They've moved on.
How ethical would it be to sell things to these grifters, to sell the shovels they will use? But I'm always hung up on the idea that they will use those shovels on others and exploit them.
Example asserts include basic string checks, regex, is-json, cosine similarity, etc. (and LLM self-eval is an option if you'd like).
I particularly like promptfoo's support for CI, which I haven't seen anywhere else, and is very important for developers pushing prompts into production (esp since OpenAI keeps updating their models every few months...).
Regarding randomness, the initialization of the weights is random, and if they use dropouts that is random too, plus the order in which to process the text might be random.
In the 90s and even 00s, the theory first crowd was mainstream and the empirical first crowd was considered fringe. Very fringe.
Personally I appreciated it when LeCun was like: “you can’t find the solution if you only search where the lamplight is shining.” Or other early deep learning practitioners note that ML theory is usually so far disconnected from practice in terms of tightness and bounds that you might as well ignore pure theory completely.
Anyway, it wasn’t until deep learning methods really smashed benchmarks across the board did people give in to the black magic / alchemy driven approaches of empiricism based upon intuition and bias developed through long-held experience.
But literally the first sentence of the readme is "Prompt engineering is kind of like alchemy."
That's empiricism. Scientific method implies formulating falsifiable theoretical claims, producing meaningful analysis of experimental conditions, publishing peer-reviewed and reproducible results.
You're trying to correct people on a subject you know nothing about.
https://www.nationalgeographic.com/magazine/article/leeches-...
Science would be replicable, ie demonstrate that this particular approach to prompt writing yields better prompts than some baseline prompt writing approach across an array of different problems.
But these models are changing literally every day, so there's no fixed thing to reveal.
So no one will be able to reproduce anything at all.
This makes all this "engineering" pretty ridiculous in my eyes, it's literally for one models bizarre emergent properties.
Although, even software engineering being an exact science, it is a funny one: most of us don’t get certified as like, let’s say, mechanical engineers do. Would they say we are engineers?
So perhaps the “engineer” term got overloaded in recent years?
If we are being pedantic and literal, this is exact in the sense that for identical seeds you get identical results.
I wound hate having to studying for a C based test when the area I work in is all web tech. Same could be said of Java. I learned it in school and haven’t used it in 10 years except for it being the backend on my first front end dev project.
I've always assumed there's a near 100% overlap between people using the term wrongly to describe any programming activity, and people complaining that it has no meaning or is self-aggrandisement
People seem to think that “automation will create new jobs”, but in the age or AI, those job opprtunities will be very temporary as the companies making the AI automate that thin layer.
Similarly, people think that humans will control AIs. That’s a bit quaint, a bit like humans controlling a corporation. The thin layer of “control” can be easily swapped out and present an improvement in the market, so that the number of totally autonomous (no human in the loop) workflows will grow.
That can include predictive policing with Palantir (thanks Peter Thiel!), autonomous killbots in war etc. Seeing how reckless companies have been in releasing the current AI in an arms race, I don’t see how they would be restrained in a literal arms race of slaughterbot swarms and panopticon camera meshes.
PS: I remember this exact phase when computers like Deep Blue beat Garry Kasparov. For a while he and others advocated “centaurs” — humans collaborating with computers. But the last decade hardly anyone will claim that a system with humans in the loop can beat a system that’s fully automated: https://en.m.wikipedia.org/wiki/Advanced_chess
If coders can call themselves engineers, no reason why anybody else solving puzzles for a living can't.
One day I think there will be true software engineering. When that happens you won't be able to start software projects without certifications, and most people (or programs!) who actually do the coding will be following careful plans and instructions from the engineers who designed the project.
I for one am very happy software isn't fully professionalized yet!
Sounds really bad.
Software can't fall on your head and kill you, not all of it at least.
Different software should require different professionals building it.
And it's usually not about the software but about the management telling the engineers to take shortcuts or whatever (Boeing comes to mind)
I recommend reading the entire postmortem, http://sunnyday.mit.edu/papers/therac.pdf yes it is quite long but if you write code in any capacity it’s worth the read
I cannot use the word software engineer, since its nothing like real engineering.
Real engineering was def harder, more math intense, and the stakes were sooo much higher. While many software problems can cause you to lose money, engineering problems can cause you to lose time. Yes it sucks your CAD designer had everything on a 0.05 degree angle and it costs 1M to redo the tool, but it also costs 16 weeks to redo the tool. We'd even offer to pay absurd money, prevent future business, etc... to get the tool done in 8 weeks, but its impossible to get it done faster. Now everything in the company is 8 weeks behind schedule.
Anyway, real engineering was harder, but programming pays soooo much more money. Its a demand thing, not a difficulty thing.
It's always been a demand thing, not a difficulty thing.
Assembly and safety critical C can be considered software engineering.
Anything with abstraction, no.
But I just mean to say that software engineering itself is definitely still a real thing and there are many people out there that can and should call themselves software engineers.
"The engineering design process, also known as the engineering method, is a common series of steps that engineers use in creating functional products and processes. The process is highly iterative - parts of the process often need to be repeated many times before another can be entered - though the part(s) that get iterated and the number of such cycles in any given project may vary.
It is a decision making process (often iterative) in which the basic sciences, mathematics, and engineering sciences are applied to convert resources optimally to meet a stated objective. Among the fundamental elements of the design process are the establishment of objectives and criteria, synthesis, analysis, construction, testing and evaluation.[1]"
It is not dependent on the problem domain, rather on how the work is performed.
I'm guessing you were trying to say something else here. I literally cannot think of a single software engineering problem I've ever encountered that didn't cost time. By your definition, then, software engineering is engineering. Your claim and your definitions are at odds with one-another.
Also, you don't directly claim it, but you seem to imply, that software engineering can't have real-world consequences... or something? As another reply points out, sometimes software is in the critical path for things like rockets and airplanes, where mistakes cost lives.
And some people making software for less life-altering systems take their craft just as seriously. Some people think that losing $10M every single second while their software is failing is a big deal.
Are you claiming people who write HFT code, ad arbitrage code, code that powers the front page of Apple, Amazon, Microsoft, and Google are just cowboying it through the day, doing nothing special?
Overall I just find this comment very confused. Maybe you could put some thought into what you're trying to say, and say it better?
Knot engineer, floor polishing engineer, water spillage cleanup engineer, tax minimization engineer, heart engineer, bus operating engineer, flower pruning engineer
Whilst BASIC/JavaScript/etc are all magic incantations to a child, a child will soon figure out there's underlaying logic, and learn the ability to reason about what code does, and what certain changes will do.
With prompts, it's all faerie logic. There is nothing to learn, there are only magic incantations that change drastically if the model is updated.
Worse yet, the incantations cannot be composed. E.g. take the SQL statement "SELECT column FROM table WHERE column = [%s]". For any given string you insert here, the output is predictable. You can even know which characters would trigger an injection attack.
With prompts you cannot predict results. Any word, phrase, or sequence of characters may upset the faeries and cause the model to misbehave in who knows what way. No processing of user-input will stop injection attacks.
Whilst it's dubious to call current software development practices "engineering", it's utterly ridiculous to do so for prompt-writing.
However, even with formally composable languages like JavaScript, a semblance of unpredictability — akin to the "faerie logic" metaphor — still persists. Languages evolve over time; Python, for instance, with its various imports that constantly disrupt my code, serves as a good example. This is perhaps the reason behind the emergence of containers to ensure code consistency.
While some elements may be more "composable" than others, it appears increasingly unrealistic in today's world to encapsulate thought processes or interactions with systems within a rigid logical framework. Large Language Models (LLMs) will keep evolving and improving, making continual interaction with them unavoidable. The notion that we can pass a set of code or words through them once and expect a flawless result is simply illogical.
I firmly believe that any effective system should incorporate a robust user interaction component, regardless of the specific task or problem at hand.
even with formally composable languages like JavaScript, a semblance of unpredictability — akin to the "faerie logic" metaphor — still persists
And they're ridiculed for it, and as you state, we design around them or replace such systems entirely.
making continual interaction with them unavoidable
Technology is never unavoidable or "inevitable". We can choose not to use it, or when to use it.
The notion that we can pass a set of code or words through them once and expect a flawless result is simply illogical.
Yet that is what we expect when we put these systems into production use, especially when many proposed use cases are user-facing and subject to injection attacks.
Whether it be the writing of adcopy, the processing of loan applications, or generating code, mistakes in these tasks have very real consequences.
Reminds me of raising kids...
Sure, the results are not deterministic in that 100% of the time the exact prompt returns the exact same result, but you can tune your prompts so that 100% of the time they give you a valid result in the result category you were seeking, and with a specific probability distribution of available choices.
Prompts are functions that can take concrete input and create a probabilistic output that can be automated upon. Especially if you only need to output one token, i.e a number, boolean, word, object reference. And for obvious reasons - the further you forecast out in a sequence the less accurate you will be.
As long as you don't change the underlying model, in a massive model with billions of parameters, there are definitely mechanisms and behaviors to discover that you can reason about.
You can't though, that's the issue. Illustrative here are tokens like "SolidGoldMagikarp", but this does happen to "normal" sequences of tokens as well.
There is no filter you can build to keep out such mistakes, any set of otherwise normal tokens could trigger the model to produce wrong output.
Because of how large these models and most prompts are, even slight changes in things like attention can cascade into extremely different results.
there are definitely mechanisms and behaviors to discover that you can reason about.
It's faerie logic. The behaviours are mere trends and observations, not underlaying truth.
The faeries reward you for offering them fruit. But offer them apple which fell from the tree exactly 74 hours ago down to the second and they'll kill you. There is no way to know ahead of time which things will upset them.
The risk here is that you're fooled into believing these systems are understandable, that you know how they work, and that you'll mistakenly use them for something where the wrong results have consequences. You'll stop double-checking the output, all humans are lazy like that, and then you'll have disaster on your hands.
Why do you think rockets explode, bridges collapse, etc.
We need to move away from prompt-engineering - it's AI-Management. You pretend you're speaking to another (albeit confusing/confused) person when extracting work from a model. You're coaxing things out of it based on hearsay and mysticism that work most of the time. Sounds a lot like AGILE and free pizza to get a junior to stay late and deliver on time.
That's not engineering, that's management.
> The word engineer (Latin ingeniator) is derived from the Latin words ingeniare ("to contrive, devise") and ingenium ("cleverness").
we, the IT crowd, are long over-due for this formalization of the professions.
The idea of an engineer that researches, designs, tests, and measures and a programmer that implements seemed to cost too much (not just monetarily) for the industry that employees them and there isn’t the sufficient need to regulate all sun-groups of the industries that employ software programmers / engineers.
> find customers who didn't rent a movie in the last 12 months but rented a movie in the 12 months before that
GPT-4 solves this without a problem [2]. Combining logic like (without additional database schema added):
> find all users who lives in Paris using lat/lng and who visited the south of France within the last month
GPT-3.5 can't understand this at all, GPT-4 solves it [3].
[1]: https://www.postgresqltutorial.com/postgresql-getting-starte...
[2]: https://aihelperbot.com/snippets/cljy8km2h0000my0fgq8kut5w
[3]: https://aihelperbot.com/snippets/cljy8q6gz000al70fvfzxt2hh
That's pretty much how I see every nightclub, to be fair...
I loved illegal raves when I was in college. Paying big bucks to listen to loops triggered by some influencer feels like a surreal parody...
Pain and joke too often come as closely bundled as within Adams's pointed work.
Just ask ChatGPT to not say a thing.
Recently I’ve been trying to engineer a prompt that I intend to run 1k times.
Noticing GPT4 bug out on several responses, I’ve talked it through the problem more and asked it to rewrite the prompt. So an automated approach to help build better prompts based upon held out gold data is useful to me.
> Your job is to rank the quality of two outputs generated by different prompts. The prompts are used to generate a response for a given task. You will be provided with the task description, the test prompt, and two generations - one for each system prompt. Rank the generations in order of quality. If Generation A is better, respond with 'A'. If Generation B is better, respond with 'B'. Remember, to be considered 'better', a generation must not just be good, it must be noticeably superior to the other. Also, keep in mind that you are a very harsh critic. Only rank a generation as better if it truly impresses you more than the other. Respond with your ranking, and nothing else. Be fair and unbiased in your judgement.
Source: https://github.com/mshumer/gpt-prompt-engineer/blob/main/gpt...
Unless your prompt seriously conflicts with the schema, it's pretty consistent.
There should be at a minimum a way to print which comparisons were made, so that you could double-check if you agree with those.
Also, we could modify the prompt to explain why the decision was made.
I was hoping this would do something more interesting with multiple messages (if using a chat model) rather than just dumping the entire prompt in one message. The assistant lets you do stuff with examples.
My first thought when looking over this tool was "Why do I have to do all the work?", the ideal scenario is that I give the high level description and the LLM does the hard work to create the best prompt.
Very nice idea, thanks. i might steal that.
https://www.reddit.com/r/ChatGPT/comments/12cvx9l/compressio...
$("[data-type='ipynb']").style.width = '100%'it reminds me of those aircraft that folks in rural india build from time to time.
> Your job is to rank the quality of two outputs generated by different prompts. The prompts are used to generate a response for a given task.
> You will be provided with the task description, the test prompt, and two generations - one for each system prompt.
> Rank the generations in order of quality. If Generation A is better, respond with 'A'. If Generation B is better, respond with 'B'.
> Remember, to be considered 'better', a generation must not just be good, it must be noticeably superior to the other.
> Also, keep in mind that you are a very harsh critic. Only rank a generation as better if it truly impresses you more than the other.
> Respond with your ranking, and nothing else. Be fair and unbiased in your judgement.
So what factors make the "quality" of one prompt "better" than another?
How "impressive" it is to an LLM? What even impresses an LLM? I thought as an AI language model, it lacks human emotional reactions or whatever.
Quality is subjective. Even accuracy is subjective. What needs testing is alignment-- with your interests. The thing is hardcoded to rate based on what aligns with model hosts' interests, not yours.
Only the "classification version" looks capable of making any kind of assertion:
> 'prompt': 'I had a great day!', 'output': 'true' [sentiment analysis I assume?]
The rest of the test prompts aren't even complete sentences, they're half-thoughts you'd expect to hear Peter Gregory mutter to himself:
> 'prompt': 'Launching a new line of eco-friendly clothing' [ok, and?]
The one for 'Why a vegan diet is beneficial for your health' makes some sense at least, but it's really ambiguous.
I'm just some idiot, but if I were creating this, I'd expect the response to ask for a number of expected keywords or something to measure how close each model comes to what the user actually wants. Like, for me, 'what are operating systems' "must" mention all keywords Linux, Windows, and iOS, and "should" mention any of Unix, Symbian, PalmOS, etc.
All tests should tank the score if it detects fourth-wall-breaking "As an AI language model/I don't feel comfortable" crap anywhere in the response. National Geographic got outed on that one the other day.
lololoollool
Many fields of study have this fuzzy property - it's easier to name which fields don't.
It's a response to the notion that the quote above comparing writing prompts to alchemy is so wrong that it's funny in some way. I thought it was a pretty good analogy.