If that’s what you want to call it.
I see hypemen overpromising and product underdelivering. And when pressed about specifics, attempts to drown queries in jargon or an ass-covering retreat to treating it like it’s just a tech demo not intended to be used for anything ever.
And it seems likely that with an order of magnitude better hardware it might be good at some things that it seems really bad at now. So yes, it's a tech demo for a lot of things that aren't quite ready, and it's also very useful as it is.
i work at a large consulting firm everyone knows and am seeing this first hand. I'm not on the bandwagon until i see real money in the bank from large AI projects succeeding. It's not happening yet. I not a naysayer but am still very skeptical.
I resent some LLM implementations on principle, but decided to give these code helpers a try. What I found was they’re reasonably bad, and I kept telling them the solution doesn’t work, only to be presented with a little tweak.
So I don’t see the point of outsourcing my thinking, I’d rather remain intelligent and do the search/try/tweak on my own, instead of pretending a half-assed LLM is genius.
That doesn’t mean they don’t have good use cases, or aren’t an improvement on previous tech. But we definitely should stop calling them mind-blowing. Jaron Lanier had long ago predicted we’d willingly downplay human intelligence to pretend AI was… I.
You seem to be doing what GP is pointing out.
GP's claim is that the relative improvement itself is mind-blowing, not that the tech is mind-blowing in an absolute sense.
I tend to agree: much of the detraction hangs on current-state rather than a probable potential-state informed by recent relative advancements.
In other words, many proponents are chuffed because of that potential; not because actuality. Likewise, skeptics are reserved because of the actuality, and not because of potential, for whatever reason (there are at least a few main ones, I think).
I feel like people are taking Moore's Law, which is definitely a real thing, and thinking that everything else is going to advance like semiconductors did, and I just am not seeing it in any other field. It's not true in software development (where gains, such as they are, are more linear than exponential) its definitely not true in rockets or civil engineering or steel or anything like that. So I am afraid a whole lot of people are expecting Moore's Law type improvements in AI, when really AI advances more like punctuated equilibrium: a sudden dramatic improvement, then a long period of consolidation and stasis, then another sudden dramatic improvement, often in a totally different unpredictable area.
But I've just been keeping tabs on AI since the hot way to do it was Expert Systems back in the 1990's, and I'm aware of its history since Norbert Weiner wrote Cybernetics back in 1947, and this seems to be a repeating pattern: a single major breakthrough (in this case, honestly, the combination of large quantities of data with NN's- with driving and natural language being the two easiest to get, and so the most prominent examples) followed by a lengthy fallow period where not much appreciable progress happens, then another breakthrough, often orthogonal to where earlier breakthroughs happened.
Edit: I haven't even started on the social and political impact this has either.
Yes if you tell it to write an entire program it will get it wrong and you'll spend some time verifying things. But that's not a sane way to use it. As a very clever auto-complete it's fantastic. It's also pretty great at getting past "blank page syndrome". Even if what it spits out is wrong it's still helpful to get you started.
Do you really think it's fine blowing 400 watts because you can't be arsed to think or do not have the creative intelligence to get over the blank page syndrome and have to lean on a crutch?
Yes, I think it is absolutely 100% fine.
From my past history:
"Is "192.168.1.4" included in the subnet "192.168.0.0/16"?" -> No
"Check whether a widget overflows in flutter" -> returns a function that cannot be made to works even with a lot of massaging (uses stuff that does not exist)
"Write a parser for this multiline format in C++ (describe format)" -> parser only read first line
Admittedly a trick one: "Can you give me a C++ function to merge 2 uint32_t and one uint16_t into a unique uint64_t?" -> happily gives an answer
Sometimes it is salvageable, and sometimes it can provide ways I did not consider to solve a problem (though the proposed solution is usually broken), but usually I would have been faster to do it myself than to try to fix whatever it gives me.
I have basically given up on it, except for some generic "how would you solve problem X?", and when I see people talking about it, it feels like a totally different world.
> "Is "192.168.1.4" included in the subnet "192.168.0.0/16"?" -> No
ChatGPT is not good at numbers or complex maths like this.
> Check whether a widget overflows in flutter
I mean this would be closed as unclear on StackOverflow, but again, this is basically asking ChatGPT to write an entire function. It can do a good stab but it's not going to get it correct.
Copilot isn't for that sort of thing. Let me give you a more realistic autocomplete example from my code:
std::fs::write(&sv_path, sv).expect("error writing top.sv");
std::fs::write(&rs_path, rs).expect("error writing top.rs");
std::fs::write(&cargo_toml_path, cargo_toml).expect("error writing Cargo.toml");
std::fs::write(&cpp_path
It completes `, cpp).expect("error writing main.cpp");` which is actually exactly what I had. I may have used Copilot to write that; I don't remember. The point is it is 100% correct and saved me writing all that. Traditional autocomplete can't compete with that.However even for "do it all for me" queries it can still be useful. For example I asked:
> I have a C++ process paused in a debugger (lldb). It is consuming a lot of memory. Is there any way I can see what is using the memory? E.g. a heap profiler that can attach to the process?
It got it wrong and told me to run the process under Valgrind. I said:
> Those instructions aren't for attaching to an existing process.
And it told me about the Heaptrack project with can do exactly what I want. I can see why search engines are integrating this asap. Even though it hallucinates, it does that infrequently enough that it's still really useful. I mean it hallucinates frequently but even so it is very useful.
Perhaps a better example is this:
> Write a C++ program to run a child process and send data to its stdin and stdout while it is running.
After 2 more prompts I got it to output the code below. This code is not correct. But the fixes to make it work are relatively minor, and EASILY less work than writing it all from scratch.
In a few years we will look on people that don't use Copilot (or similar) like people that don't use IDEs. They're hurting their own productivity out of principal.
-------
#include <iostream>
#include <unistd.h>
#include <sys/types.h>
#include <sys/wait.h>
#include <vector>
class ChildProcess {
public:
ChildProcess() : pid(-1), pipefd{-1, -1} {}
~ChildProcess() {
if (pid != -1) {
close(pipefd[1]); // Close write end of the pipe
waitpid(pid, nullptr, 0);
}
}
bool create(const std::string& command, const std::vector<std::string>& arguments) {
if (pipe(pipefd) == -1) {
std::cerr << "Failed to create pipe." << std::endl;
return false;
}
pid = fork();
if (pid < 0) {
std::cerr << "Failed to fork process." << std::endl;
return false;
} else if (pid == 0) {
// Child process
close(pipefd[0]); // Close unused read end of the pipe
// Redirect stdin and stdout to the pipe
if (dup2(pipefd[1], STDIN_FILENO) == -1) {
std::cerr << "Failed to redirect stdin." << std::endl;
return false;
}
if (dup2(pipefd[1], STDOUT_FILENO) == -1) {
std::cerr << "Failed to redirect stdout." << std::endl;
return false;
}
// Convert arguments to a C-style array
std::vector<char*> args;
args.reserve(arguments.size() + 2);
args.push_back(const_cast<char*>(command.c_str()));
for (const std::string& arg : arguments) {
args.push_back(const_cast<char*>(arg.c_str()));
}
args.push_back(nullptr);
// Execute the child process
execvp(command.c_str(), args.data());
// execvp() only returns if there's an error
std::cerr << "Failed to execute child process." << std::endl;
return false;
} else {
// Parent process
close(pipefd[1]); // Close unused write end of the pipe
}
return true;
}
void write(const std::string& data) {
if (pid != -1) {
::write(pipefd[1], data.c_str(), data.size());
}
}
std::string read(size_t numBytes) {
std::string output;
if (pid != -1) {
char buffer[numBytes + 1];
ssize_t bytesRead = ::read(pipefd[0], buffer, numBytes);
if (bytesRead > 0) {
buffer[bytesRead] = '\0';
output = buffer;
}
}
return output;
}
std::string readLine() {
std::string output;
if (pid != -1) {
char buffer;
ssize_t bytesRead;
while ((bytesRead = ::read(pipefd[0], &buffer, 1)) > 0) {
output.push_back(buffer);
if (buffer == '\n') {
break;
}
}
}
return output;
}
private:
pid_t pid;
int pipefd[2];
};
int main() {
ChildProcess childProcess;
std::vector<std::string> arguments = {"arg1", "arg2"};
if (childProcess.create("child_process", arguments)) {
childProcess.write("Hello, child process!");
std::string output = childProcess.read(1024);
std::cout << "Child process output: " << output << std::endl;
std::string line = childProcess.readLine();
std::cout << "Child process line: " << line << std::endl;
}
return 0;
}One thing I always find funny is the general expectation that machine learning models are both incredibly generalised and designed based on the way biological systems work, but should also be 100% perfect and never be wrong just like a machine and NOT like a biological system, those things are mutually exclusive; even the best, smartest most physically capable humans will still sometimes spill their coffee, yet we expect coffee-bot 2024 not to do this.
Certainly machines can be much better at a task than humans are, but if that tasks requires generalisation then it's still gonna fuck up from time to time.
Machine learning programs may think a dog is actually a cat sometimes, but afaik they ain't ever called their teacher "Mum" yet.
Aren't you overgeneralising a bit? Not even I would say that CNNs for image classification, or Deep-RL for board game-playing are a "continual disappointment" and they certainly predate LLMs. Are you talking about NLP? Even Neural Turing Machines were quite capable in language pairs with large parallel corpora (and similar linguistic structure).
Basically, what do you mean by "continual disappointment"? What I'm aware of is an incessant hype crescendo that crashing over everything like a relentless wave.
RLHF trained GPT3 and then DALLE/Midjourney/Stable Diffusion changed all that. Suddenly AI not only got good, but the field broke loose of the inane and insane obsession with pseudo-safety that had been holding it back. Now the rest of us can use it without dropping $5M on a GPU cluster and hiring a dozen researchers first. AI is no longer a disappointment.