Especially since most desktop applications people use are web apps. Of the native apps people use that leverage this sort of stuff, almost all are GPU accelerated already (eg. image and video editing AI tools)
Especially since most desktop applications people use are web apps. Of the native apps people use that leverage this sort of stuff, almost all are GPU accelerated already (eg. image and video editing AI tools)
Maybe your local machine can run, I don't know, a model to make suggestions as you're editing a Google Doc, which frees up the Big Machine in the Sky to do other things.
As this becomes more technically feasible, it reduces the effective cost of inference for a new service provider, since you, the client, are now running their code.
The Jevons paradox might kick in, causing more and more uses of LLMs for use cases that were too expensive before.
I like the latter. Why even use a new word if it's just going to be the same as "client"?