AI Prompt Engineering Is Dead
spectrum.ieee.org
spectrum.ieee.org
It's medieval anti-enlightenment black magic. The genies are black box tangled confections of brute force computing power in reach of only nation states or megacorps, which also don't know how they work but try to weave them along their agendas.
Programming was supposed to be a conquest of reason over the darkness of ignorance, and this is the exact opposite. :D We're back to wizard whispering to spirits. "Prompt engineer" a circle of salt and a burnt white rabbit to invoke Eschmoûn for a good harvest.
What a depressing thing to aspire toward.
An electrician is a magician already. A painter is a sorcerer. A programmer creates life.
I think it's cool I can understand an algorithm and what it does, but, once it becomes machine code, my understanding utterly ends. What's the CPU doing with it? Have I even delidded my CPU to make sure there's actually silicon in there and not pixie dust? Forget the computer; how does my own brain work? Nobody even knows how brains store memories. A neurologist might know a lot more than I do, but maybe this is all a solipsistic simulation where I'm a brain in a jar and the neurologist is just an NPC controlled by a fancy AI.
Perhaps, in reality, the feeling that you know what's going on was always a foolish conceit. It's magic all the way down. Astronaut with gun: always has been.
I guess I have to find comfort in degrees because otherwise I feel very helpless and stupefied by that way to view the world.
Not that ChatGPT can't do it, just that a lot of people seem to think it's going to be like talking to the star trek computer.
So many desk jobs that the professional managerial class does today can probably be described similarly. That doesn't mean that they can't still be high-paying.
The word "just" is doing a LOT of work there!
[ and it's already too late for this, but before posting snarky notes about prompt engineering as a profession you should know that this article isn't about that - it's about tools like DSPy that automatically optimize prompts ]
It's the same story with search engines. Google et al may like to say they've tuned the engine to be "optimal" at guessing what people "really want." In practice, I've found the opposite to be the case. The damn thing interprets my query in such a liberal way that it returns a bunch of irrelevant garbage. Maybe their numbers show that this works better for the average person. But with people who really know what they're looking for (like me) it seems to have gotten far worse.
Prompt engineering is still important.
From a philosophical point, it always makes a difference how you phrase a problem. If you formalize a problem in Lean code, it's easy to understand for a proof assistant but not for most humans (even mathematicians or programmers).
So rephrasing the problem in human language makes it easier to grasp and comprehend. And given the dataset the AI is trained on, there might be certain style it prefers. If you train it on Lean code, better input lean code. Also, more context makes it usually easier to solve the problem.
You need to do a good job at it, but it’s not something anyone is going to get a full time job to do.
To do that, they'd need to have some kind of coherent detailed plans that actually made sense in the context of all the minutia of the rest of the system. If the product were built of fingerpaint and good intentions, POs would still need programmers.
Prompt engineering is kinda like a basic form of data science. You have a dataset and some manually labelled results, and you hypothesize on prompts and try to improve on some metric. You'd be surprised how the tiniest of changes will alter a result. Thinks like a hyphen, semicolon, or capitalization can sway the metric. It's very finicky and very annoying.
e.g. why does the JSON output have silly whitespace/quotation sometimes? Obviously it's because the first token the model output was `{` and not `{"` like it should be. Obviously.
GPT-4-turbo's output context window is limited to 4096 so fixing this is relevant. You can use logit_bias for it.
You also need to know the tools to use and there's a lot of settings you play with that are hidden from the user in mainstream apps.
Well, That would be ideal, but if I type in "Middle-aged white male in full plate armor standing on a battlefield resting on a full tower shield" I likely will want to further modify the result, or style, or detail level. There almost certainly will continue to be "hacks" to get it stylized as desired. Even if I say "Painting of..." there's still a huge range of options.
I understand and agree that it's desirable to get AI prompts as close to natural language, but how do you quantify a level of stylization in natural language? "A very very very very Michelangelo style painting of a slightly slightly slightly Middle-aged white male..."
I think prompt engineering will change quickly, and to keep up, it could potentially be a 'profession' that is very specific to the model. I don't think that's a bad thing, but I would think/agree that it will likely not employ many people at all.
The product I’m building is predicated on prompt engineering and it works really well for us — our results need to be as accurate and structured as possible so they can be re-used as software dependencies. We do this by auto-generating a context on the fly and tailoring it to your specific data sources (databases, APIs), sampling data, and so on. We do both manual and "auto" tuning of the context to provide the best results.
Even if the system adds a layer with "autotuned prompts" as the article calls them, some prompts will inevitably work better than others, and some models will have to be "spoken to" in a certain way or use certain keywords to make it work in a specific way, but that knowledge may not work with all models.
For example, recently I've discovered that adding "no yappin" at the end of the prompt makes ChatGPT 4 go directly to the interesting part, but if I use it on another model it could do nothing, or maybe it gets confused and starts spouting garbage, or it may get pissy and scold me for being rude.
This might come off as a joke, but I don't doubt at all that it will soon be reality, as different models will have different "personalities". This term is not to be taken literally, but it gets close enough to the idea of representing the internal workings of a LLM.
Little tricks and hacks, like going as far as bribing or threatening the LLM to get a better or specific kind of output, mostly model-specific.
The future will be interesting for sure.
If correct, there are a few relevant conclusions that you can draw. The first is that prompt engineering is a real thing. By talking "correctly" to the LLM you're providing it with additional information that will allow it to give you more legitimate answers. This might help explain these threads where I'm seeing a lot of comments from people who rave non-stop about all of the work they're able to get LLMs do for them AND also people who cannot get LLMs to do anything worth doing (I'm included in this latter group). The former know how to "talk right" to the LLM.
While I think prompt engineering is real, I don't think it's necessarily important. "Talking correctly" is nebulous and resists our ability to categorize. The metric is "gets good answers from LLM", but in order to know if the answers are good you have to also know the reality. It's pretty obvious (at least to me) that if you have to choose between speaking in a useful way and knowing the reality that you would just choose knowing reality[3].
The other thing is that this means that LLMs are a dead-ish end. Because they're approximating the intelligence in natural language, they'll never be smarter than our collective words. Sure, we'll be able to make cheaper, smaller, and more specialized LLMs, but they'll need another non-LLM component to reach the next level of usefulness.
[1] - And I mostly keep mentioning it to see how people react. I've yet to see a "this is definitely wrong because ..." type of response. I'm hoping to see one that actually has some meat to it.
[2] - So, the idea is that word2vec allows you to do vector math of the pattern: KING - QUEEN = v; Man - V = Woman. Great, so our grammar has some sort of algebraic structures inside of it. And this sort of makes sense that this would happen. Kind of an information theory version of frequent messages should be small. Only here, it's natural language should have a structure that mirrors the structure of problems we care about. People who talk the right way have a tendency to solve problems that matter and natural language evolves in the direction that enables people to be successful.
The problem then is that some problems are simply too complex to embed their structure into natural language grammar. Like rocket science, brain surgery, or OS construction. The other problem is that some problems are simply too niche to develop a jargon. And finally, some problems have a structure that we haven't seen before. In any of these cases, I expect LLMs to fail to be useful.
[3] - I suppose there could exist some world in which the knowers are able to certify the talkers and then maybe that's easier than teaching unknowers to be knowers. But it feels like it's a bit risky. If the talkers ever get off base, then it's not like anyone is going to notice until it's too late.