My worry is that these personal assistants will all be provided "free" by giant companies and hosted in the cloud so I won't want to use them.
My worry is that these personal assistants will all be provided "free" by giant companies and hosted in the cloud so I won't want to use them.
The more worrying fact to me, is not MegaCorp will be providing it, but that it's possible that self-hosting would be technologically infeasible. That is, even if MegaCorp gives you a free and open copy of personal assistant, it's still not practical run it on your own hardware - massive computing power and storage is required to operate it, and massive data is needed to train (and continue to train) the machine learning model. Only MegaCorp can afford to run it, and provide it as a service.
This has already happened in web search, content recommendation, speech recognition, machine translation, and many more applications. It's simply impossible to run your own speech recognition or machine translation package that has a comparable quality to Google's.
And these technologies will eventually become an integral part of life in a future society - e.g. When you must use machine translation or personal assistant in your daily life, and when only MegaCorp can provide a high-quality solution, it's a recipe of disaster. It's Ghost in the Shell.
Unfortunately, I'm definitely seeing it as a very likely outcome.
The only possible way to stop it, is to continue advancing computing power according to Moore's Law - which is, fortunately, not entirely impossible if alternative computing technologies are developed in the future. But still, there's still the ML training problem.
No, economies of scale will always result in some services being cheaper to centralize.
I think the same will happen with 3D printing and solar panels. No matter how cheap they get, it's always even cheaper for the power company to build a huge solar farm, or for me to order a printed part online and have it shipped. (For small batches anyway)
Services like search and speech recognition and translation that benefit from having large datasets to work with, also need fast Internet connections to keep that data updated.
So in 2020 my $100 might buy as much CPU as Google's $80, but it only buys as much network transfer as Google's $10.
The ratio is probably worse if you consider that self-hosted or even federated search means all the web scraping has to be replicated to 100x or 1,000x more databases than Google has to manage even with their CDNs. Redundancy is not free.
I want it to work too, but I don't think centralization is a technical problem with a technical solution. We're fighting uphill against a law of economics. Not even capitalism, either, economics itself as a law of nature.
> services being cheaper to centralize.
You are talking about its relative cost, while I'm talking about the mere possibility.
Back in the early 80s, having a personal workstation (not a home computer) to process private information was nearly an impossibility, not simply because it's cheaper to centralize but because it's technologically prohibitive. Even in the early 90s, I still read stories about how a hacker in a crypto community needed to import his private key every time he comes to work in the morning to read some personal messages, and to remove it again at the end of the day. The server could not be trusted and there was no alternative.
And today, everyone uses Gmail and the same privacy problem is still here - the server can't be trusted, but at least, if you are personally willing to spend some money and energy on your project, it's possible to host your own. It'll never be mainstream, yes, but it's possible and Moore's Law enabled it.
Maching translation, or personal assistant, on the other hand, now resembles a "personal-workstation-in-the-80s" problem. If computing power continues to advance according to Moore's Law, at least it's going to be a "Gmail-in-the-20s" problem after a decade or two.
> We're fighting uphill against a law of economics.
Moore's Law and personal computing, regardless of how sufficient or insufficient, is historically a balancing force. If this force continues to act upon, the level of centralization of future software applications will still be imaginable to us.
But if it ceases to exist, things will become much worse.
That is all I wanted to say.
I don't really agree with that. All of these things can be done on your own machine, and generally comparably or even in a superior fashion to Google. Bear in mind that Google has to focus on a very long tail of "serving every single user", whereas your software has to serve you.
This means your speech recognition needs to support your speech, not the speech of every accent and dialect on earth. Your content recommendation can be limited to types of content that are even peripherally of interest to you, and likely in a limited subset of languages.
Basically any amount of data and processing Google needs to do something for everyone, you can massively subset for just yourself. The biggest issue is that most of the data and code that companies like Google are using is proprietary, so good luck getting a copy to run.
Just a quick response here:
> This means your speech recognition needs to support your speech, not the speech of every accent and dialect on earth.
What about machine translation? I don't need French, German, or Russian translation, until the day I need it... Hypothetically, combining it with speech recognition, it will create a perfect platform for wiretapping conversations.
My worry is built on your worry; I worry that people won't even think of using their self control to say no to this, but will wait breathless for some third party to save them from their own bad decisions.
That's all but guaranteed if it's any time in the next decade or so.
I always think Machine Learning is merely a useful and flexible pattern recognition machine that can be trained, anything beyond that is all media hype and marketing, there's nothing to see here.
Now my opinion has changed - if a pattern recognition machine that contains no breakthrough (GPT-3) can be surprisingly intelligent if you throw a lot of data to it, any new development cannot be worse than that. I'm now a true believer on the feasibility of AI assistants.
And that’s why I disagree with the claim that GPT-3 is in any way intelligent. It’s a chameleon. Or a Chinese room [1].
I think to have real intelligence you need a model for reality, not just a language model. Otherwise you can do nothing more than parrot and remix the words of others. Parrots are entertaining companions though.