The big question is how much better it's going to get. We may be plateauing in capability-- or at least slowing in capability growth-- or we may not be.
But it doesn't need to get better for it to eat a lot of peoples' lunch.
I think the real potential gains come when they start extending the architecture, but as long as it's just scaling up and different training/inference regimes, then it seems it'll be more of the same rather than game changer.
Given how connected the whole SF/AI scene appears to be (alchohol + drugs too?), it's hard to imagine a company the size of OpenAI not leaking. If there had been any amazing discoveries there, I think there'd at least be rumors (and not just "they seem to have something called Q* going on").
For instance: Chemists using the AI backwards. Take some phrase like 'vanadium increases the Young's modulus versus silicon in 1040 steel' and then work the AI backwards from that phrase. As in, assume the AI outputted that phrase and see what inputs were most likely to generate it. There may be some real discoveries just by working an AI backwards iteratively.
Simple things like that are still open and attainable right now, it's just that AI is still so young that we really haven't explored all of what they do yet.
Right. Regardless of any "superhuman" abilities, we can talk to our computers now, and they can talk back. The decades old holy grail of HCI has been achieved. That fact alone is going to change everything.
LLMs sound so sure of themselves and people think "well i'm dealign with the most advanced technology ever so it must be right...".
On the implementation side of things, it's hard for me to get the non-deterministic aspect of llms right in my head. I put an LLM and RAG system in prod with a team and went through rounds of the usual testing. 99 times it passed but on test 100 it would fail, so you'd adjust the system prompt. Then it'd pass 500 times and fail at 501. Adjust the system prompt, then it would pass 9 times and fail at 10. That system went to production but there's the low level worry in my mind, when is it going to fail to give the correct output? The fact that you can never guarantee the output of an LLM from a given input severely limits where they should be used IMO. I don't think it's wise to have the output of an LLM be the input to another program, there's no functional relationship between domain and range with an LLM.
That problem is usually met with "well, a human would make the same mistake.." but the reason computers exist is to do long, tedious, lists of tasks/instructions very fast that humans get wrong. Simulating a human, and all those imperfections, with digital logic seems contradictory to me.
edit: Also, just want to point out that the "testing" mentioned in my post was all manually done by humans. You can't automate testing the response of an llm unless you use another model to grade the response as correct or not but then you're right back to not being able to trust that the grader will always act consistently.
1. At present, AI can be massively overhyped, especially by salespeople who have an incentive to overhype it.
2. Current generative AI has made leaps and bounds improvements in the past few years, and it's present capabilities provide an enormous productivity tool for those who use it intelligently.
I mean, no, I don't think GPT 4 is AGI, nor do I really think it's that close. That said, every time I use it I'm amazed at how uncannily good it is, and it is able to save me a ton of time.