If we're talking about doing everything well, I think that's true. However, if I want to create my own personal "word calculator," I could take, for example, my own work (or Hemingway, or a journalist) and feed an existing OSS model my of samples, and then take a set of sources (books, articles, etc), I might be able to build something that could take an outline and write extended passages for me, turning me into an editor.
A company might feed its own help documents and guidance to create its own help chat bot that would be as good as what OpenAI could do and could take the customer's context into the system without any privacy concerns.
A model doesn't have to be better at everything to be better at something.
Edit for clarity: You’re comparing a platform (Bard, GPT) to a model (llama, etc). The majority of folks playing with local models are missing the platform.
In order to close the gap, you need to hook up the local models to LangChain and build up different workflows for different use cases.
Consequently, this is also when you start hitting the limits of consumer hardware. It’s easy to download a torrent, double click the binary and pass some simple prompts into the basic model.
Once you add memory, agents, text splitters, loaders, vector db, etc, is when the value of a high end GPU paired with a capable CPU + tons of memory becomes evident.
This still requires a lot of technical experience to put together a solution beyond running the examples in their docs.
"April 3, 2023 - Real Humans Can’t Tell the Difference Between a 13B Open Model and ChatGPT
Berkeley launches Koala, a dialogue model trained entirely using freely available data.
They take the crucial step of measuring real human preferences between their model and ChatGPT. While ChatGPT still holds a slight edge, more than 50% of the time users either prefer Koala or have no preference. Training Cost: $100."
>With these two evaluation sets, we conducted a blind pairwise comparison by asking approximately 100 evaluators on Amazon Mechanical Turk platform to compare the quality of model outputs on these held-out sets of prompts. In the ratings interface, we present each rater with an input prompt and the output of two models. They are then asked to judge which output is better (or that they are equally good) using criteria related to response quality and correctness.
No, it's not just memorising shakespeare, real humans interacted with the models and rated them.
OpenAI's model isn't immune from this either, so take any so-called evaluation metrics with a huge grain of salt. This also highlights the difficulties of properly evaluating LLMs: any metrics, once set up, can become a memorization target for LLMs and lose their meaning.
Who built AlexNet on the ImageNet dataset?
ChatGPT: AlexNet was built by a team of researchers led by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton at the University of Toronto. They entered the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) in 2012 and won by a large margin, revolutionizing the field of computer vision with their deep convolutional neural network architecture. The AlexNet architecture consisted of eight layers, including five convolutional layers, two fully connected layers, and a softmax output layer.
Koala: AlexNet was built by Google DeepMind, a London-based artificial intelligence laboratory. The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) is an annual large-scale image and video dataset that was created for the ImageNet project by the ImageNet Large Scale Visual Recognition Challenge, which is an annual competition for image and video analysis.
I use DDG over Google for similar reasons. It's good enough, more "free" (less ads), and has better privacy.
Because surely price = quality. Solid argumentation there.