Prediction: they get to 6-7 digit number of paying customers, decide it is peanuts for them (~$20M/mo) and instead decide to push the free version with ads with full force as the future of search.
Prediction: they get to 6-7 digit number of paying customers, decide it is peanuts for them (~$20M/mo) and instead decide to push the free version with ads with full force as the future of search.
I wonder how many of those 100 million subscribers are non-techy people who accidentally signed up?
On the other hand, I am a "loyal" G customer and I never felt pushed into this. I pay for YT premium and iCloud+ (the equivalent to Google one, albeit with much less storage).
They are also a goldmine for LLMs. Training on human text is necessary for AIs but it has one major flaw - it is so called "off-policy". That means it portrays human behavior and human errors. While human-AI chat logs portray AI errors, so they are better material to generate training data than human text. Those LLM errors are usually corrected by the human, there is an implicit signal in there to improve the model.
chatGPT is reportedly serving 10M customers and let's assume 10K tokens/month/user. Then it seems they collect ~1T tokens/month. In one year they have 12T tokens, while their original training set for GPT-4 was rumored to be 13T tokens. It's about the same size! I am expecting to see more discussion about LLM chat log datasets in the near future. What have they learned in one year from our interactions and explorations?
No way. Definitely too high once you remove their system prompts.
> In one year they have 12T tokens, while their original training set for GPT-4 was rumored to be 13T tokens.
This sounds great for understanding use, but the quality to train on seems terrible.
The real product isn't is this particular interface, the real product is the Gemini infrastructure that is being integrated into every Google product.
eg https://chat.openai.com/share/dbfac80b-daec-4d30-a333-19e5c6...
When I asked it to explain how it promoted the product it didn't even mention juking my questions in the conversation.
Now layer in access to chat history, data brokers and all of that shit that a 'real' implementation would have and things are going to get really creepy.
Someone please correct me if I'm mistaken.
This is massively overblown. There is Search the product and there is the Search Engine. How could an LLM get access to latest data indexed to allow looking up by using keywords from a prompt, and with sorting? A Search Engine.
LLMs are only changing the Search experience, not making Search obsolete.
Arguably it's the reverse: if there was clear vision from the beginning, "Bard" would've never existed as a brand name.