It can also simulate a Zizek vs Wittgenstein argument over Russian literature. the fact that it can usually write executing computer code is nothing short of magic to me.
What one would want is something like LLaMA2-Coding-Vue2, maybe with a LoRA for the library or concept.
I wrote a big word salad about all the other things you could do but suffice it to say there's no reason anyone should be limited to a single pass on a monolithic without automatic context window augmentation nor automatic code checking & regeneration & model escalation (e.g. query out to a 200B coding model or something).
I'm javascript centric these days, but the same concept should work for most languages with a package manager.
Some of these features could even be integrated into say yarn/npm directly, you currently have a devDependencies section of a package.json..I could imaging having something along the lines of a "llmDependencies" section to define which main model and version to use and which "library" models to use.
If you don't pay for ChatGPT, you get GPT-3.5. You can also get access to GPT-4 if you use the playground.
We do a lot of experiments involving gpt3.5, 4, claude-v2, titan-large, and palm2, and for what it's worth, on our real production workloads gpt4 shines. We can make Palm2 produce decent results with a lot of extra effort, and claude-v2 is passable but gpt4 does not disappoint. This is low-grade knowledge management stuff, and we are not using it as a information-retrieval system - but for basic 'cognitive' tasks where all the information needed is provided in the prompt. I'd not rely on it for info retrieval tasks such as the examples quoted above - its knowledge base is highly compressed, after all.
When there's one definitive answer to something that people keep repeating there's a slight chance that it's actually true. Shocking, I know.
But hey, if you are really looking to convince yourself of something, I have no doubt that it can be done.
I keep going around telling people that 1+1=3, why do they always give me the same nonsense about the number '2'?
I blame Sam Altman.
2. Obnoxious? This is my opinion, I don't get what Sam Altman or 1+1=3 is supposed to mean in this context.
I asked it (actual names changed):
"I run the Linux command line program "foo". When I use the flags -xyz, I get results, but when I use -txyz I get nothing. What could this mean?"
And it told me: "The lack of results is because you didn't use the -t flag".
Or I ask it some very basic music theory questions and it gets stuff wrong all the time, giving impossible answers.
Do you have API access? the old model there still gives me very good results.
I'd kill to have an easy, wont-get-me-banned-way to submit a query to both the UI and API at the same time and show the results in, say, meld or so.
Tools have limitations.
How do you get that out of an LLM? What tool is any good if it doesn’t work 100% of the time predictably?
Query many times on the right model(s) for the question and the correct answer will be there 99.999% of the time as the other hallucinations will be thrown out.
It’s funny those the LLM haters keep raising the bar to a level that no other software can reach. Chatgpt is a tool like any other and just like a hammer, it can be misused, or it can be incredibly useful if used well. I personally find chatgpt mind boggling and astounding and use it every day multiple times for both coding and non coding purposes. But it’s totally normal and reasonable to me to expect bugs, do you really expect visual studio to run on large solutions and never have crashes or memory issues or slowness? If so you’re going to be disappointed.
I find that no examples often leads to the same result as a befuddled junior, but with examples often it gains confidence. Also I find playground to give me much better code snippets than chat sometimes
There's a lot of "extraneous" information and details that appear useless and unrelated, but that long-time developers have tucked away in their brains (or really anyone who has done something at a high level for a long time), that turns out to be incredibly useful; generally these people don't require examples - they just know what the right answer is, because they've been exposed to a variety of problems over a long career or lifespan.
That's where LLMs need to be to be truly "useful". I think if we get them to that point, we'll really have something useful on our hands.
Both young and old shall feel the pain then.
I think the future is extremely well-trained base models that are then fine-tuned on specific domain knowledge. I'm already seeing that with Meta's Llama 2 personally at home. My company has access to Azure OpenAI Service's GPT-4 trainable model that we've been working with on all our documentation, and the results look very promising so far.
But, this idea of the All-Knowing Oracle that can answer any question you have is all well and good, but it's so far beyond our current processing power as to be a pointless endeavor at the moment. Will we get there eventually? Yeah, sure. According to Jim Keller as of around 2019/2020 when he was speaking on Lex Fridman's podcast, he said we still have room to go 1,000,000 times smaller - meaning chips. I think we probably need two more orders of magnitude in processing power of the strongest high-end GPUs before we're there. NVIDIA's H100 is a great achievement, but we need something about 100x as fast, and I think we'll still need dozens of them working in parallel to train the model that can answer all these questions.
All of this though is a moot point, because it's the lawyers that have already slowed down progress. OpenAI is terrified of being sued. Altman couches it behind terms like, "AI Safety", "AI Alignment", etc., but it's fear. It's all stemming from fear. And it's all stemming from people just not "getting it".
We're entering a new age of upheaval, and there's going to be rogue AIs that tell people to go kill themselves to reduce climate change. You know why? Because there are humans that tell people to go kill themselves to reduce climate change. These models are language models. We taught them how to think, and we taught them how to think like us, so it's no surprise to me whatsoever that they behave like us - meaning they occasionally lie and they occasionally go off the rails and go a little crazy.
We have become gods and we have made a creation in our own image. Most of the time it's awesome, sometimes it's a little wild and wacky.
However a few times there have been some mechanical refactoring-style grunt work I've delighted to have been able to let ChatGPT do. However, the rate ChatGPT is giving me subtly wrong results is just high enough that I end up cross-checking everything, and then it takes me a bit more time than it would've otherwise taken. Give it a year or two, maybe?
Maybe? It’s not clear what would bring a qualitative improvement, barring massive amounts of new training data.