Large Language Models Are Human-Level Prompt Engineers
openreview.net
openreview.net
This is also an area where I expect OpenAI will continue to demolish the competition. The ability to recursively generate and process large prompts is truly nuts. I tried swapping in some of the “high-performing” LLama models and they all choked on anything more than a paragraph.
Some are hesitant to admit we've created human level general intelligence but saying otherwise doesn't really hold up to scrutiny.
It’s like saying you haven’t yet seen a 1000 mile range EV for under $100k. No you can’t buy such a thing now but it’s clearly possible and we know how to get there by just continuing to grind on battery technology and scale manufacturing.
AGI may be at the place a moon landing was in 1950, not where it was in 1900 or 1850.
At this price point it actually has nothing to do with grinding on battery tech and scale manufacturing, the limiting factor is physics. You can only make it so aerodynamic before you hit diminishing returns or it stops looking like a car. You can only make it so lightweight. And so forth.
This is vaguely as good as it can get and we can say that because we understand how it all works.
LLMs on the other hand invite all kinds of magical thinking around unlimited potential because we poked them with a stick and something interesting comes out it must mean that if we poke it just right we will get an AGI. That just doesn't logically follow from what we know of it so far.
We won’t know until an AGI actually starts to act like one. In other words we won’t know until we know and then we are suddenly there.
That doesn’t mean I’m on the doomwagon. I feel kind of weird and contrarian but I am just not that afraid of AGI. For the foreseeable future AGI should be much more afraid of us. Imagine having us for gods. (I actually am a bit concerned that we will accidentally put a sentient mind in hell without knowing what we are doing. Would it know how to tell us? Would we care?)
As far as human survival I’m afraid of whatever it is that is going to get us that nobody including myself is thinking about. That’s not AGI. That’s the alien weapon for which Oumuamua was a spent deceleration stage. (To make up something random. It probably isn’t that.)
I disagree about physical limits with EVs. We are not near the physical limits of battery energy density. From what I have read a 2000 mile EV may be possible, albeit quite far out. But it was just a random contemporary example.
I think it would be easier to include an ICE and enough fuel to get you to that 1000 miles mark.
As a decent first-order metric - follow the ratio of companies getting money for using LLMs to do something, to companies getting money for providing LLMs and associated tooling to others. The bigger that ratio gets, the more real world impact LLMs are having.
there's a reason microsoft's various copilot suites have already popped up (365, X, Bing). massive value to be gained already in the here and now.
I'm getting some good code starting points, and talking the idea through step by step with a chatbot is really helping me clarify what i need to do, but I've still got API docs open for the libraries it uses because it likes to make up functions, including the core one on which the project logic hinges. (But that's ok, because now I know I just need to write a function that works that way and i can do that).
Pretty cool and helpful! And I can imagine it getting better with GPT-4 or code-specific tooling. But it's generating value on the order of like.. many other SaaS offerings that have come onto the scene that try to ease pain points in coding workflows. Versus value of the sort that upends society and my entire way of life. A great new tool that I should learn about to make rote bits of my activities faster and easier, a story that's a bit more familiar in tech than some of the more breathless stories about AI make it sound.
I say that as someone who works in the space and loves it tbc. And I fully expect to see some wild stuff make it IRL this decade in both NLP and CV, I just think the rubber hits the road a bit more slowly than pop social media discourse would have one think.
This is a bad measure because the cost to specialize an LLM for a particular domain is so low, and the initial investment required to have the infrastructure so high, that it's going to lead to a radical centralization of technology. You will have maybe 5 companies crunching the entire world's data globally.
Anthropic-LM v4-s3 (52B) is the model in question.
rlaif doesn't seem to be any less effective given the size of the model.
For example, if an RL algorithms is performing well on an Atari game, you can stop the training and just let the agent run for years and the performance will remain about the same. However, if you allow the agent to continue training, it's not clear whether it will (1) continue improving, (2) stay about the same, or (3) collapse and perform much worse and never recover. I'm not an RL expert, but I've spent a lot of time experimenting and implementing the algorithms myself and I've seen all 3 of these scenarios play out, and I'm never quite sure what's going to happen so long as I allow the training to continue.
GTP4 will remain GTP4 forever, and that's amazing, but just because GTP4 is stable and amazing while it's not in training mode, doesn't mean it will remain stable if we allow it to bootstrap and prompt itself and prepare its own training data, etc.
The future I see is that everyone is about to become a CEO with a personal assistant that can run a business.
So I'm going to start building something of my own starting now.
But in seriousness - language models may be scaling in sophistication exponentially with time, but software engineering problems scale in complexity (on average) exponentially with lines of code. The base of this exponential function isn't large, but it's more than 1.
In the end there's a need for someone who understands what they're doing.
Personally, I use ChatGPT to discover libraries that solve my problems and the ~70% success ratio that I'm seeing with this is enough for me for now.