I agree that it's fundamentally different, but I'm not exactly sure how, and I think it's subtler than you're suggesting.
I agree that it's fundamentally different, but I'm not exactly sure how, and I think it's subtler than you're suggesting.
We call computers deterministic despite the fact that they don't with perfect reliability perform the calculations we set them. The probability that they'll be correct is very high, but it's not 1. So the requirement we have for something to be considered deterministic is certainly not "perfectly a hundred percent of the time", as the parent to my comment suggested.
It's a non-deterministic algorithm, of which many kinds exist. Producing different answers that are close-ish to correct is in fact what a Monte Carlo algorithm does. Not that you'd use GPT3 as a Monte Carlo algorithm though, but it's not that different.
Close-ish to correct makes sense for some problems and makes no sense at all for others.
Imagine a C compiler that does aggressive optimizations - sacrificing huge amounts of memory for speed. On one hand, it even reduces computational complexity, on the other it produces incorrect results for many cases.
GPT-3 as presented here would be comparable to that. Neither are equivalent to executing the original code.
Meanwhile, the result of something like gcc is, even if it runs on a computer with faulty RAM.
Speed and memory is orthogonal to my point, which is about the output of two methods of arriving at an answer. I'm obviously not saying GPT-3 is anything like as efficient as running a small function.
What distinction are you drawing between the output of an interpreted program and a compiled program?