>We can only speculate what data closed LLMs were trained on, but I'd be highly surprised if Google/openai had exclusive access to a bigger repository of written data than, well, the internet, as it presents itself to the world at large.
>What exactly would they require "user data" for?
There are several classes here:
A) Total internet data. Google/OpenAI may have more data from Google Books/GSuite/etc. but maybe not. No way to know. Maybe even if they do, it's not significant compared to total data volume. Since we can't meaningfully compare, let's just ignore it.
B) Global usage data. This is useful to further tune the model - we saw what the open models could do with a partial log of ChatGPT. OpenAI of course has the entire log. For example, it's possible that users in country X ask for stuff in a different manner, or that terms have a local meaning the model may not be aware of. Language evolves after all. A local model can at most update on current user data, or by much slower updates from the origin, and OSS has less resources here.
C) Local usage data. For example, a company may wish an LLM to access all its documents to create a knowledge base. There's a good chance all the documents stored in Office 365/GSuite. You can guess who has easy access and who gets the scary permission prompts. Another example: The LLM writing an email may wish to be aware of the previous communication in the thread and your general tone. Or replace Spotlight/Windows search with an LLM, but the LLM needs access to all your data to properly search it. Some of this can be emulated with really long prompts, but it's more efficient to just let the LLM have access.
>Even the existing cloud based solutions don't need access to user-data either to perform their functions.
Currently no, but the personal assistant they want to build will require it.
>Products can be developed by basically every group with the passion to do so
>By doing what, deploying ever larger models?
Alas, OSS devs tend to get bored on 'non-sexy' subjects. Meanwhile, Microsoft and Google will embed LLM in all their apps. The apps have their own moats (data migration, UX) and in turn act act as a moat for the LLM.
Moat in action:
Imagine Thunderbird worked with OSS LLM and Outlook works with OpenAI GPT. A user has meetings in Outlook and uses GPT to do various related planning. Say the user was willing to migrate to OSS LLM. But OSS LLM doesn't have easy interface with Outlook (Microsoft 'competitive' behaviour), and manually importing all the time is too messy. The user may even consider switching to Thunderbird, but Thunderbird doesn't do ActiveSync, and IT refuses to even consider allowing IMAP in its Exchange, so user is stuck with Outlook and in turn with OpenAI GPT. Doing ActiveSync is boring for OSS devs, so Microsoft gets an indirect moat: Exchange <=> Outlook <=> OpenAI.
SD is way more in tune, subject is way more popular with devs I guess. These people have a chance. They don't however need to brag about inevitable victory of SD or how Adobe is going down.
>> Everything needs to be deployed, and BigCorp can just push it as an OS update.
>And app providers can just update an app.
Deploying an app requires more friction. How do you get users to install it in the first place? Not impossible (see Google Chrome over Internet Explorer) but a struggle where the OS maker has a built-in advantage.
>By doing what, deploying ever larger models?
A bit of that, but I expect more effective tuning because they have way more usage data.
>Attention based transformers have O(n^2) scaling
There are numerous papers trying to improve that. We'll see.