The investments are absurdly large but if you believe LLMs will be fundamentally transformative then I can understand the logic.
The big context window is so nice, I can dump in huge files or conversation exports and it can handle them without problems
Am I the same only one who uses the web interface?
I keep thing simple (KISS) because I teach for a living, and only things my student can pick up in a class are useful to me.
I have a slide in a presentation on the topic that's an inverted pyramid, as pretty much the entire LLM field rests on only a few companies, who aren't necessarily even doing everything correctly.
The fact that they even have a seat at such a small table with so much in the pot already and much more forecasted means they command a large buy in regardless of their tech. They don't ever need to be in the lead, they just need to maintain their seat at the table and they'll still have money being thrown their way.
The threat of missing the next wave or letting a competitor gain exclusive access is too high at this point.
Of course, FOMO driving investments is also the very well known pattern of a bubble, and we may have a bit of a generative AI bubble among the top firms where large sums of money are going to go down the drain on investing into overvalued promises because the cost of missing a promise that will actually come to fruition is considered too high.
Ironically the real payoff is probably in focusing on integration layers at this point, particularly given the gains in performance over the past year in research by developing improved interfacing with SotA models.
LLMs at a foundational pretrained layer are in a race towards parity. Having access to model A isn't going to be much more interesting than having access to model B. But if you have a plug and play intermediate product that can hook into either model A or B and deliver improved results to direct access to either - that's where the money is going to be for cloud providers in the next 18 months.
The big context window is pretty magical for some use cases. There are lots of things RAG with limited context can't do.
Having a large context window is pointless unless the model is able to attend to attention on the document submitted. As RAG is basically search, this helps set attention regardless of the model's context window size.
Stuffing random thing into a prompt doesn't improve things, so RAG is always required, even if it's just loading the history of the conversation into the window.
Assuming a new unheard of play called "Romeo and Juliet " RAG can answer "Who is Rosaline?".
But a large context window can answer "How does Shakespeare use foreshadowing to build tension and anticipation throughout the play?" "How are gender roles and societal expectations depicted in the play" or "What does the play suggest about the power of love to overcome adversity?".
In other words, RAG doesn't help you answer questions that pertain to the whole document and aren't keyword driven retrieval queries. It's a pretty big limitation if you aren't just looking for a specific fact.
RAG limits you to answers where the information can be contained in your chunk size.
I dunno I kind of get the perception Microsoft's locked up OpenAI's funding so Google couldn't throw money at them even if they wanted to?
https://docs.anthropic.com/claude/docs/give-claude-room-to-t...
You can see how the AI arrived at the response.
Claude is more creative and that comes with a higher rate of hallucinations. I hope I'm not misremembering but GPT-4 was also initially more prone to hallucinations but also more capable in some ways.