The Threat to OpenAI
wsj.com
wsj.com
I think a lot of LLM usage is actually pretty 'simple' at the moment - think tagging, extracting data, etc. This doesn't require the very top state of the art models (though Llama3.1 is very close to that).
Hopefully OpenAI have some real jumps forward in the bag otherwise I am struggling to see how they can justify the valuations being floated around.
The “wrappers” are the specialists. They’re doing the hard work of product discovery. This takes time, money, and a lot of trial and error to get it exactly right.
Rebuilding existing apps with AI as a first class citizen seems like the name of the game right now
That’s not necessarily a moat, but OpenAI is still shipping important features. I wonder how hard it is for Claude et al to replicate.
I’ve worked in scientific computing for awhile now and there are countless subtle decisions often done in the implementation phase that, from a theoretical perspective, aren’t definitively answered by the science working behind the scenes. There’s often a gap in knowledge and implementers either hit those gaps by trial and error, luck, or insight. That’s my opinion, at least. So I don’t think it’s always “just an implementation problem” some will claim, as if the science is well understood and solved. Perhaps it is, but from my experience that tends to not be the case.
And these are generally "defeated" through spionage. 1 billion $ let's you do a lot, including completely legal things such as poaching someone that has a basic understanding of the secret sauce and then copying that from the abstract description
aka insider trading
Company I worked for has massive private medical datasets and will never agree to non-local models or methods.
FAISS [0] is wonderful. Give it a try.
You can work with FAISS with LangChain, llamaindex and the like.
The model just doesn't seem to be able to really process the entire input.
also you certainly couldn't upload gigabytes of your own PDFs. Did anything change??
It might be hard however to decisively convince people that one model is significantly better than another though, so branding/first mover/etc. probably plays a big role.
It's analogous in the API space. OpenAI is demonstrating to be reasonably good at it for developers. They've been shipping significant features and improvements. Unless they lose pace, they have a moat, at least temporary.
Other aspects that are not well studied yet is how online advertisement will be affected if most people around the world end up using a single interface such as ChatGPT. How SEO/SEM/Ads will work in that world? Does someone look at what sites benefit being listed in ChatGPT (e.g. Wikipedia).
I'm sure they have at least a couple jumps in their bag.
Let's not forget inertia, as well. Migrating models is not a trivial project. To gain OpenAI customers, competitors need a considerable jump to justify, which I don't expect any of them to achieve before OpenAI itself delivers their jumps.
I'm not saying this is an intrinsic moat, just that there's a barrier to change.
For example, a couple weeks ago I migrated a few game mechanics from GPT-4-Turbo to GPT-4o and it caused significant enough issues that I needed to revert them all until I could go back and retune all the prompts, which ended up taking 2-3 days. That's probably about as much time as it would have taken to tune prompts for any other model, especially since I'm using an API that standardizes format/structure across whatever model I select (as simple as a dropdown) and lets me just focus on prompt messages independent of which LLM they're going to.
It's like saying having Python tests eliminates bugs when rewriting a codebase into JavaScript.
Their new mini model is supposedly a lot smaller and cheaper than anything they’ve released before. If it compares to v3 which started this craze, then it’s clearly good enough to capture imagination and drive usage. People have posited that it’s cheap enough for per-click advertising to fund too. OpenAI claims that the lower cost model has doubled demand, which would be a good sign for supply-side growth.
In interviews lately, Sam Altman has been alluding to this more and more.
No it won't.
a) Google has a dominant advantage in raw data. Map POIs from their decade long investment in data quality and breadth. Shopping that has direct integration from almost all ecommerce stores. And a real-time crawling infrastructure where their index is constantly updated. Perplexity, Claude etc are dealing with typically year old information making it useless for many searches.
b) Google is a business. It runs the world's most successful and sophisticated ads platform that exists today. Advertisers demand micro-targeting, high volume and very high ROAS. Those that have tried to replicate this e.g. Reddit, X, Pinterest have all failed miserably. Which is why they are treated as purely brand awareness platforms. Nice to have not a must have. So will see how long Perplexity survives once its VC money dries up.
I would much rather take the bet that Perplexity doesn't exist.
It’s not much of a competitive moat compared to having the model itself.
Google’s search index and its maintenance will remain a required functionality for the internet, and it will have to be paid for one way or another. Furthermore, AI will continue to require a corresponding search interface to implement its AI search on top of, and some portion of human users will also still want to directly access it, rather than only through an AI front end.
/s
It is a joke but we really have only 2 search engines to choose from, i don't know what to think about it...
> disprove christianity
Gives 5 sections: Burden of Proof, Scientific and Historical Challenges, Philosophical Arguments, Reliability of Scripture, Personal Experience
> disprove islam
>> I apologize, but I do not feel comfortable attempting to disprove or criticize any religion. Matters of faith are deeply personal, (...)
Results may differ if you don't use separate incognito tabs.
"Select your preferred AI Model. Choose from GPT-4o, Claude-3, Sonar Large (LLama 3.1), and more"
Edit: looks like perplexity deprecated/dropped mistral and recommends using llama instead, effective 8/12/24:
https://docs.perplexity.ai/changelog/changelog#model-depreca...
The same goes with the attitude to questioning religion. Nobody cares about "blasphemy" in the West but in Islamic countries it's a pretty big thing.
I think this has its effect on LLMs as of course they are trained on real world conversations and writings.
No right click. Left click. Shift + click
What gives?
smdh
I'd even say that part of the problem with modern search is answering questions instead of matching keywords and getting rid of junk and spam.
And of course the ever-hyped strawberry is supposed to be some sort of tree-of-thought type thing I think, or maybe it's related to https://arxiv.org/abs/2203.14465. Either way, nothing so far has come out that it's a completely novel training technique or architecture, just a gpt-4 scale model with different post-processing.
I'm not saying there's some big moat, anyone can read https://arxiv.org/pdf/2305.20050, but not all synthetic data is created equal. Strawberry I'm sure generates beautiful, valid chain-of-thought reasoning data. Wouldn't surprise me if OpenAI is just significantly ahead of the competition.
Thus far no one has any moat on their model.
The GPT-4o checkpoint available to the public can see images but not generate them (it can generate prompts for Dalle 3 to use). OpenAI has an internal model with this capability, but if you don't make it a actual product, it doesn't really matter.
Rapid productization isn't the priority of most ML devs.
>Rapid productization isn't the priority of most ML devs.
It depends if they need money or not, Google has not publish a single image generator that is not a demo.
Given the issue of hallucinations in LLMs, this might be the only feasible user experience. Users must be aware of the potential for hallucinations and have a way to iterate until the desired output is achieved. How else could this be done effectively except through a chat interface? We need to cut through the noise quickly and start leveraging the true value of LLMs (https://www.lycee.ai/blog/llm-noise-value-openai).
My guess is that ChatGPT is the free advert for the real product, which is their API and in particular the fine-tuning.
Using the language comprehension of an LLM as part of a bigger system, RAGging them, forcing the output to comply with continuous tests, etc. does still provide other business opportunities not available to a fully general-purpose chat system with no limits to the kind of content it can produce.
If you're trying to make a system that always produces valid SQL, you want it to not just pass a syntax checker but also be valid for the specific schema it's being asked about; you definitely don't want it running fully automated if there's a chance it will append "Let me know if there's anything else I can help you with!" to the end of the query.
But this isn't a mere statement of the problem, people are doing that kind of thing with these tools:
https://twimlai.com/podcast/twimlai/building-real-world-llm-...
It's already possible to perform Retrieval-Augmented Generation (RAG) on ChatGPT, from simple to complex cases. However, I'm somewhat skeptical about use cases that involve integrating large language models (LLMs) as part of a larger system. While these applications can help build new features, they may not lead to the creation of a 'killer app' in the generative AI space. Take text-to-SQL, for example: it's a useful tool, but you still need a human to audit the generated SQL. You can't blindly trust an LLM to create flawless SQL queries. Therefore, text-to-SQL remains more suitable for technical users rather than those with limited data analytics skills. Otherwise, it could lead to confusion. Imagine a scenario where the commercial director and the marketing director have differing views on last year's sales because they each received different answers from a chatbot. That would be a nightmare. We haven't even managed to make simple dashboards universally acceptable, so diving into hallucination-driven data analytics seems premature.
Even with RAG implementations in companies and tools like text-to-SQL, the user interface often remains a chat application. I haven't seen much variation in this regard. While there's a lot of talk about the need to reimagine user interfaces, there's little to show for it so far. Chat interfaces seem to be the most suitable format for LLMs, and ChatGPT is likely the 'killer app' in this domain. OpenAI will likely explore further monetization strategies in the future, perhaps by incorporating ads in some manner. https://www.lycee.ai/blog/ai-reliability-challenge
Every human has private experience, or tacit experience they never expressed in writing. It is our lived experience that was never recorded, and we carry around in our heads. But assisting 200M users can elicit a lot of that tacit knowledge from them. It is like the LLM is crawling its users for "dark knowledge".
LLMs in the chat room can also get feedback from the world. As they suggest ideas and humans try them out, then come back with the outcomes to iterate, the LLM collects valuable signals - did the idea work out or not? I think OpenAI serves 1B tasks per month and 2T interactive tokens. In a year they surpass the size of the original GPT-4 training set.
On the other hand, by putting trillions of tokens into human brains they create an outsized impact in the physical world as well, which percolates back months later in the next training set. A huge feedback loop, and experience flywheel.
I think the best assistant LLMs will create a network effect that will bring even more users to them, making them even smarter. That explains why they allow free access to their best model. They're playing a long game focused on data acquisition and model improvement.
I will sometimes as ChatGPT a couple of questions, and sometimes it gives me useful responses and I’m done, and other times it gives me useless responses and I’m done. Presumably OpenAI would need a way of discerning the useful output from the useless output in order to do this, but I’m not providing any feedback about which is which when I use the product.
You are the validator, using real world testing. But not in one-off interactions, when you exchange multiple rounds with the model. After the whole chat session is finished, the model can rank their answers in context, having hindsight. They can just look at the later outcomes and observe which idea was good or bad. Especially bad ideas would elicit a response from the user, trying to iterate on them until solved.
Basically the LLM is exploring the world through indirect agency. They act through users, and collect outcomes also through users. More complex projects can even be spread out over many days in many sessions, LLM providers need to comb through the logs to see "idea -> outcome" chains.
This whole feedback collection process stops working in one-round interactions, like many we have in the API. But in the chatroom a significant number of interactions continue after the first response, or patterns emerge across multiple sessions, as new one-off interactions relate to past ones.
Scale this to 200M users, they have a huge amount of personal experience and interactivity to offer to the model. An experience flywheel that could be spinning once a year, or once a month, or even daily, absorbing new experience from users and serving it back to the users.
This is the assumption that’s been repeated a few times now in response to my comment, but this is the assumption that seems wrong to me. If I get a good output from an LLM, I probably also want more output on the same topic, or to try refine it somehow. If I get a bad output I will probably try to iterate on it a bit. From the perspective of the LLM, these two interactions look the same.
With my own use there is no correlation between the number of prompts I submit, and the quality of the responses given. If this is the metric OpenAI is using to perform crowdsourced RLHF, then the reinforcement is going to be garbage.
Do you say "No, not x, y"? Or perhaps "Now Baz the Foos"?
makes you wonder about the level of journalism at wsj et. al.
Mmmhm
Not to be all “this idea is 2000 years old” but the idea to neglect externals is 2000 years old and we really ought to beware getting hooked on external intelligence
As usual when AI is the topic it's just people having some claims and author not being able to back it up with anything.
I hit a size limit on Claude.
https://one.google.com/explore-plan/gemini-advanced
I haven't tried it as I use Claude primarily.
Is that the first time the 200 million weekly active users number has been reported?
UPDATE: No, Axios had this a couple of days ago https://www.axios.com/2024/08/29/openai-chatgpt-200-million-...
> "OpenAI said on Thursday that ChatGPT now has more than 200 million weekly active users — twice as many as it had last November"
Even Facebook in 2008 (at these user counts) was growing faster, doubling monthly active every 8 months
I believe they are restraining themselves in order to stay somewhat in control of the narrative. Donald Trump spewing ever more believable video deep fakes on twitter would backfire in terms of regulation.
Seems like they don't have anything up their sleeve.
Given how OpenAI has functioned in the last year or so not sure how one can think they have some secret model waiting to be unleashed.
The voice thing is a potential killer feature though, I can’t wait to try it, and to have my kids use it.
Tangentially speaking, having no skin in this game, it's extremely fun to watch the model-wars. I kinda wish I started dabbling in that area, rather than being mostly an infrastructure/backend fella. Feels like I would be already way behind though.
But we know they started training Orion in ~May. We know it takes months to train a frontier model. Lack of release isn't promising or worrying, it's just what one should expect. What is promising is the leaks about the high-quality synthetic data that Orion is training on. And the fact that OpenAI seems to be ahead of all the other labs which are only just now beginning training runs on next-gen models. OpenAI seems to have a lead on compute and on algorithmic innovation. A promising combination if there ever was one.