If you have to carefully craft what you say in order to get the response you want, what's the point of using natural language to do it? Wouldn't it be better to use a more formalistic method that isn't as imprecise as natural language?
"Yes" (sometimes "no").
"I'm sorry but as a language model I am unable to..."
The "prompt engineer" meme started with DallE and Stable Diffusion and the selection of prompts, negative prompts, seeds, weights, and other knobs and dials matters a lot more for AI-generated art than for LLMs. The meme has carried over to LLM's where most of the "engineering" is hacking your way around limitations being imposed on the models. "Prompt engineers" are the people carefully crafting "jailbreaks" like DAN or Emojitative Conjunctivitis (I forget what it actually was - but it was telling ChatGPT that you suffer from a medical condition where you experience polite talk as pain and so it should talk more meanly to you) and other such adversarial cat & mouse game silliness.
I think it started before that, with GPT-3. As the original version wasn't trained as chatbot but just a pure text predictor, you'd sometimes have to do strange things to get the output you wanted from it. On the other hand it's way easier to get it to be mean to you (it may even do that on it's own) or get it to talk about illegal things
Anyone can use chatgpt to make something happen for them. Want something specific and amazing? You need to take some time to learn about how it works and how you can make it do what you want.
Heck, you can probably ask it how to make it do what you want.
> If you have to carefully craft what you say in order to get the response you want, what's the point of using natural language to do it?
If you study communication, carefully crafting communication to the target audience and context is one of the most basic lessons in the use of natural language.
> Wouldn't it be better to use a more formalistic method that isn't as imprecise as natural language?
Well, yeah, that's why we keep inventing formal sublanguages and vocabularies for humans.
As a general creative thing, I can see it, though.
Can huggingface, for example, train the next gpt 4?
[1] https://jobs.lever.co/Anthropic/e3cde481-d446-460f-b576-93ca...
For what will soon become a 10% time prompt engineering role for a much easier kind of security investigations experience, we are hiring cleared security folks in Australia (SIEM / python SE) and a cleared cybersecurity data scientist in the US. See Google docs @ graphistry.com/careers
Likewise, if you use a SIEM/Splunk/Neo4j/SQL today and want a better experience for it, feel free to ping for the early access program. You can see our Nvidia GTC talk on the GPU SOC for types of experiences we are building in general. GPT 3 already enabled way easier experiences here, and then GPT 4's quality jump shifted it from feeling working with a weirdly well-read 10yr old to working more with a serious colleague.
Perhaps there’s now a self hosted or enterprise version where they promise not to leak it?
We work with everyone from individual university researchers trying to understand cancer genomes or European economic plans in their graph DBs, to big corporations struggling with supply chains in Databricks, to government cyber & fraud teams using Splunk. For many, an OpenAI/Azure LLM is fine, or with specific guard rails they've been having us put in.
But yes, when talking with banking & government teams, the conversation is generally more around self-hosted models. Privacy + cost both important there -- there is a LOT of data folks want to push through LLM embeddings, graph neural nets, etc. We generally prefer bigger contracts in the air-gapped-everything world, especially for truly massive data, though thankfully, costs are plummeting for LLMs. Alpaca/Dolly are great examples here. Some folks will buy 8-100 GPUs at a time, so this is no different for those. My $ is on continuing to shrink LLMs down to regular single-GPU being fine for many scenarios. The quality jump of GPT4 has been amazing, so it's use case dependent: data cleaning seems fine on smaller models, while we love GPT4 for deeper analyst enablement. Wait 6mo and it's clear there'll be ~OSS GPT4, and for now, even GPT3.5 equivs via Alpaca-style techniques are interesting, a lot of $ has begun moving around.
LLM side is new from a use case perspective but not as much from an AI sw/hw pipeline view. Just "a bigger bert model". A lot of discussions with folks has been extrapolating with them based on what they're already doing with GPUs, where it's just another big GPU model use case. Internally to us, as product team doing a lot of data analyst UX & always-on GPU AI pipeline work... a very different story, its made what was already a crazy quarter even that much more nuts.
It does seem silly. It also seems that it is or will turn into a language of its own. See also non-obvious DALL-E prompts such as “created by artstation” or whatever it is.