Do AI companies work?
benn.substack.com
benn.substack.com
We talk about scaling laws, superintelligence, AGI etc. But there is another threshold - the ability for humans to leverage super-intelligence. It's just incredibly hard to innovate on products that fully leverage superintelligence.
At some point, AI needs to connect with the real world to deliver economically valuable output. The ratelimiting step is there. Not smarter models.
In my mind, already with GPT-4, we're not generating ideas fast enough on how best to leverage it.
Getting AI to do work involves getting AI to understand what needs to be done from highly bandwidth constrained humans using mouse / keyboard / voice to communicate.
Anyone using a chatbot already has felt the frustration of "it doesn't get what I want". And also "I have to explain so much that I might as well just do it myself"
We're seeing much less of "it's making mistakes" these days.
If we have open-source models that match up to GPT-4 on AWS / Azure etc, not much point to go with players like OpenAI / Anthropic who may have even smarter models. We can't even use the dumber models fully.
Your paycheck depends on people believing the hype. Therefore, anything you say about "superintelligence" (LOL) is pretty suspect.
> Getting AI to do work involves getting AI to understand what needs to be done from highly bandwidth constrained humans using mouse / keyboard / voice to communicate.
So, what, you're going to build a model to instruct the model? And how do we instruct that model?
This is such a transparent scam, I'm embarrassed on behalf of our species.
Some examples to illustrate the point: supply chain risks and an ever increasing amount of dependencies (look at your average React project, though this applies to most stacks), overly abstracted frameworks (how many CPU cycles Spring Boot and others burn and how many hoops you have to jump through to get thigns done), patterns that mess up the DBs ability to optimize queries sometimes (EAV, OTLT, trying to create polymorphic foreign keys), inefficient data fetching (sometimes ORMs, sometimes N+1), bad security practices (committed secrets, anyone? bad usage of OAuth2 or OIDC?), overly complex tooling, especially the likes of Kubernetes when you have a DevOps team of one part time dev, overly complex application architectures where you have more more services than developers (not even teams). That's before you even get into the utter mess of long term projects that have been touched by dozens of developers over the years and the whole sector sometimes feeling like wild west, as opposed to "real engineering".
That's why articles like this ring true: http://www.stilldrinking.org/programming-sucks
However, the difference here is that I wouldn't overwhelm anyone who might give me money with rants about this stuff and would navigate around those issues and risks as best I can, to ship something useful at the end of the day. Same with having constructive discussions about any of those aspects in a circle of technical individuals, on how to make things better.
Calling the whole concept a "scam" doesn't do anyone any good, when I already derive value from the LLMs, as do many others. Look at https://www.cursor.com/ for example and consider where we might be in 10-20 years. Not AGI, but maybe good auto-complete, codegen and reasoning about entire codebases, even if they're hundreds of thousands of lines long. Tooling that would make anyone using it more productive than those who don't. Unless the funding dries up and the status quo is restored.
I think the testimonies often repeated by coders that use these code completion tools is that "it saved me X amount of time on this one problem I had, therefore it's great value". The issue is that these all fall into a research of n=1 test subjects. It's only useful information for the subject. It appears we don't realize in these moments that when we use those examples, even to ourselves, we are users reviewing a product, as opposed to validating if our workflow is not just different but objectively better.
The truth lies in the aggregate data of the quality and crucially the speed by which fixes and requirements are being implemented at scale across code bases.
Admittedly, a lot of code is being generated, so I don't think I can say everyone hates it, but until someone can do some real research on this, all we have are product reviews.
2. The original post is, to my eyes, not implying anything about training a model to help use other models, so the second part seems irrelevant.
Ehm yes. That's because it actually doesn't work as well as the hype suggests, not because it's too "high bandwidth".
Which tells me the problem may be fundamental, not a technical one. It's not just a matter of needing "more intelligence". I don't question the intelligence or skill of the people on the outsourced team I was working with. The problem was simple communication. They didn't really know or understand our business and its goals well enough to anticipate all sorts of little things, and the lack of constant social interaction of the type you typically get when everybody's a direct coworker meant we couldn't build that mind-meld over time, either. So we had to pick up the slack with massive over-specification.
Left to their own devices they often implemented requirements that were full of errors or even completely unnecessary because they did not understand the domain well enough to ask pointed questions.
Remember the old bit about the media-- the stories are always 100% infalliable except strangely in YOUR personal field of expertise.
I suspect it's something similar with AI products.
People test them with toy problems -- "Hey ChatGPT, what's the square root of 36", and then with something close to their core knowledge.
It might learn to solve a lot of the toy problems, but plenty of us are still seeing a lot of hallucinations in the "core knowledge" questions. But we see people then taking that product-- that they know isn't good in at least one vertical-- and trying to apply it to other contexts, where they may be less qualified to validate if the answer is right.
For me, the number of times where it's led me down a hallucinated, impossible, or thoroughly invalid rabbit hole have been relatively minimal when compared against the number of times when it has significantly helped. I really do think the key is in how you use them, for what types of problems/domains, and having an approach that maximizes your ability to catch issues early.
Gell-Mann amnesia:
https://en.wikipedia.org/wiki/Michael_Crichton#GellMannAmnes...
Perhaps less than before, but still making very fundamental errors. Anything involving number I'm automatically suspicious. Pretty frequently I'd get different answers for the same question (to a human).
e.g. ChatGPT will give an effective tax rate of n for some income amount. Then when asked to break down the calculation will come up with an effective tax rate of m instead. When asked how much tax is owed on that income will come up with a different number such that the effective rate is not n or m.
Until this is addressed to a sufficient degree, it seems difficult to apply to anything that involves numbers and can't be quickly verified by a human.
But. Try this approach instead: have it generate python code, with print statements before every bit of math it performs. It will write pretty good code, which you then execute to generate the actual answer.
Simpler example: paste in a paragraph of text, ask it to count the number of words. The answer will be incorrect most of the time.
Instead, ask it to out each word in the text in a numbered list and then output the word count. It will be correct almost always.
My anecdotal learning from this:
LLMs are pretty human-like in their mental abilities. I wouldn't be able to simply look at some text and give you an accurate word count. I would point my finger / cursor to every word and count up.
The solutions above are basically giving LLMs some additional techniques or tools, very similar to how a human may use a calculator, or count words.
In the products we've built, there is an AI feature that generates aggregations of spreadsheet data. We have a dual unittest & aggregator loop to generate correct values.
The first step is to generate some unittests. And in order to generate correct numerical data for unittests, we ask it to write some code with math expressions first. We interpret the expressions, and paste it back into the unittest generator - which then writes the unittests with the correct inputs / outputs.
Then the aggregation generator then generates code until the generated unittests pass completely. Then we have the code for the aggregator function that we can run against the spreadsheet.
Takes a couple of minutes, but pretty bulletproof and also generalizable to other complex math calculations.
Programming is codified applied mathematics and involves numbers in all but the most trivial programs.
> LLMs are pretty human-like in their mental abilities.
LLM's are algorithms. Algorithms do not have "mental abilities", people do. Anthropomorphizing algorithms only serves to impair objective analysis of when they are, or are not, applicable.
> In the products we've built, there is an AI feature that generates aggregations of spreadsheet data. We have a dual unittest & aggregator loop to generate correct values.
> The first step is to generate some unittests. And in order to generate correct numerical data for unittests, we ask it to write some code with math expressions first. We interpret the expressions, and paste it back into the unittest generator - which then writes the unittests with the correct inputs / outputs.
> Then the aggregation generator then generates code until the generated unittests pass completely. Then we have the code for the aggregator function that we can run against the spreadsheet.
How is this not a classic definition of overfitting[0]?
Or is the generated code intentionally specific to, and only applicable for, a single spreadsheet?
Actually this ability is at very core of human brain. Yet we humans traded most of it for the ability to speak better. Nevertheless, brain can still read very fast as well as count objects/words really fast, but as you've never really trained, this part of the brain is optimised a lot to do other stuff. Try scanning the texts diagonally, and try to come up with a random word count and what this text is meaning. On the first tries you will be making a lot of errors, but eventually, just 100 hours of training (much less when you are kid) and you will be able to scan the texts in seconds and extract the meaning and the word count very accurately. This is a real technique people use for reading fast.
The thing is – the best and most useful ability of human brain is to adapt, from the very moment it adapted to the harsh reality of being expelled from trees and living on the ground – that's how language was created in the first place, as a result of this adaptation. LLM's can't adapt today, nor doesn't need to adapt, so their mental abilities will never come close to the brain which was adapting for millions of years. Something which is not adapting will become obsolete and come to it's end, this is a fundamental law.
I actually think you would, lots of time I've been surprised by how accurate people can eye-ball stuffs.
Yes.
Suppose someone developed a way to get a reliable confidence metric out of an LLM. Given that, much more useful systems can be built.
Only high-confidence outputs can be used to initiate action. For low-confidence outputs, chain of reasoning tactics can be tried. Ask for a simpler question. Ask the LLM to divide the question into sub-questions. Ask the LLM what information it needs to answer the question, and try to get that info from a search engine. Most of the strategies humans and organizations use when they don't know something will work for LLMs. The goal is to get an all high confidence chain of reasoning.
If only they knew when they didn't know something.
There's research on this.[4] No really good results yet, but some progress. Biggest unsolved problem in computing today.
[4] https://hungleai.substack.com/p/uncertainty-confidence-and-h...
Of course, citing sources would always trump any kind of confidence rating in my mind; sources provide provenance, while confidence rating can be fudged just like a bogus answer can.
I'm not keen on marketing words like "superintelligence" but boiling it down that's what in my mind the OP said. These systems are limited in ways that we do not yet fully appreciate. They are not silver bullets for all or maybe even many problems. We need to figure out where they can be deployed for greater benefit.
Totally. They are large language models, not math models.
I think the problem is that 'some people' overhype them as universal tools to solve any problem, and to answer any question. But really, LLMs excel in generating pretty regular text.
it does 2 things. 1 tells me how deep i am in the conversation and 2 when the computation falls apart, i can assume other things in its response will be trash as well. and sometimes number 3 how well the software is working. ranges from 20$ to 35$ on average but a couple days they go deep to 45$ (4,7,9 responses in chat session)
today i learned i can knock the computation loose in a session somewhere around the 4th reply or $20 by injecting a random number in my course work and it was latching onto the number instead of computing $
Is this because it's actually making less mistakes, or is it just because most people have used it enough now to know not to bother with anything complex?
Once we have actual super intelligence there is no need for humans to innovate anymore. It is by definition better than us anyway.
I guess you could still have artisanal innovation
Computers might do some things “better” than humans, but we’re still going to do things because we _want_ to. Often the fun is in the doing.
LLM’s can vomit out code faster than me, I still enjoy writing software. They’ll vomit out a novel or a song, but I still like reading/listening to stuff a human has taken time and effort to create something, because they can.
It's a token prediction machine. We've already generated most of the ideas for it, and hardly any of them work because see below
> Getting AI to do work involves getting AI to understand what needs to be done from highly bandwidth constrained humans using mouse / keyboard / voice to communicate.
No. Gettin AI to work you need to make an AI, and not a token prediction machine which, however wonderful:
- does not understand what it is it's generating, and approaches generating code the same way it approaches generating haikus
- hallucinates and generates invalid data
> Anyone using a chatbot already has felt the frustration of "it doesn't get what I want". And also "I have to explain so much that I might as well just do it myself"
Indeed. Instead of asking why, you're wildly fantasizing about running out of ideas and pretending you can make this work through other means of communication.
The model is doing the exact same thing when it generates "correct" output as it does when it generates "incorrect" output.
"Hallucination" is a misleading term, cooked up by people who either don't understand what's going on or who want to make it sound like the fundamental problems (models aren't intelligent, can't reason, and attach zero meaning to their input or output) can be solved with enough duct tape.
But true, it's just the output.
I find this funny because for what I use ChatGPT for - asking programming questions that would otherwise go to Google/StackOverflow - I have a much better time writing queries for ChatGPT than Google, and getting useful results back.
Google will so often return StackOverflow results that are for subtly very different questions, or ill have to squint hard to figure out how to apply that answer to my problem. When using ChatGPT, i rarely have to think about how other people asked the question.
This is just another way of saying this technology, like Blockchain before it, is a solution in search of a problem.
SOTA models are impressive, as is the idea of building AGIs that do everything for us, but in the meantime there are a lot of practical applications of the open source and smaller models that are being missed out on in my opinion.
I also think business is going to struggle to adapt and existing business is at a disadvantage for deploying AI tools, after all, who wants to replace themselves and lose their salary? Its a personal incentive not to leverage AI at the corporate level.
the problem is really, can it learn "I need to turn this over to a human because it is such an edge case that there will not be an automated solution."
Makes sense.
This is the main bottle neck, in my kind. A lot of people are missing from the conversation because they don't understand AI fully. I keep getting glimpses of ideas and possibilities and chatting through a browser ain't one of them. On e we have more young people trained on this and comfortable with the tech and understanding it, and existing professionals have light bulbs go off in their heads as they try to integrate local LLMs, then real changes are going to hit hard and fast. This is just a lot to digest right now and the tech is truly exponential which makes it difficult to ideate right now. We are still enveloping the productivity boost from chatting.
I tried explaining how this stuff works to product owners and architects and that we can integrate local LLMs into existing products. Everyone shook their head and agreed. When I posted a demo in chat a few weeks later you would have thought the CEO called them on their personal phone and told them to get on this shit. My boss spent the next two weeks day and night working up a demo and presentation for his bosses. It went from zero to 100kph instantly.
I can benchmark the quality of one LLM's translation by asking another to critique it. It's not infallible, but the ability to chat with a multilingual agent is brilliant.
It's a new tool in the toolbox, one that we haven't had in our seventy years of working on computers, and we have seventy years of catchup to do working out where we can apply them.
It's also just such a radical departure from what computers are "meant" to be good at. They're bad at mathematics, forgetful, imprecise, and yet they're incredible at poetry and soft tasks.
Oh - and they are genuinely useful for studying, too. My A Level Physics contained a lot of multiple choice questions, which were specifically designed to catch people out on incorrect intuitions and had no mark scheme beyond which answer was correct. I could just give gpt-4o a photo of the practice paper and it'd tell me not just the correct answer (which I already knew), but why it was correct, and precisely where my mental model was incorrect.
Sure, I could've asked my teacher, and sometimes I did. But she's busy with twenty other students. If everyone asked for help with every little problem she'd be unable to do anything else. But LLMs have infinite patience, and no guilt for asking stupid questions!
Do you speak two or more languages? Anyone that does is wary of automated translations, especially across estranged cultures.
> It's a new tool in the toolbox, one that we haven't had in our seventy years of working on computers,
It's data analysis at scale, and reliant on scrapping what humans produced. A word processor does not need TB of eBooks to do it's job.
> They're bad at mathematics, forgetful, imprecise, and yet they're incredible at poetry and soft tasks.
Because there's no wrong or right about poetry. Would you be comfortable having LLMs managing your bank account?
> If everyone asked for help with every little problem she'd be unable to do anything else.
That would be hand-holding, not learning.
Won't that be fun.
We have a new god (superintelligence) but only the special can see it because it's too intelligent for people to interact with it.
It's so advanced all the common ways humans communicate "using mouse / keyboard / voice" don't work
The reason we've never seen it help us is we need more ideas (prayers) first.
NPC brains are really hard to understand I will say that and they use a lot of electricity.
Logic programming? AI until SQL came out. Now it's not AI.
OCR, computer algebra systems, voice recognition, checkers, machine translation, go, natural language search.
All solved, all not AI any more yet all were AI before they got solved by AI researchers.
There's even a name for it: https://en.m.wikipedia.org/wiki/AI_effect?utm_source=perplex...
> As soon as definite knowledge concerning any subject becomes possible, this subject ceases to be called philosophy, and becomes a separate science.
I'm not actually sure I agree with it, especially in light of less provable schools of science like string theory or some branches of economics, but it's a great idea.
So, what are good examples of some things that we used to call AI, which we don't call AI anymore because they work? All the examples that come to my mind (recommendation engines, etc.) do not have any real societal benefits.
It's only recently with generative AI that you see any examples of the opposite, people outside the field calling LLMs or image generation "AI".
It’s the cognitive equivalent of the Indiana Jones sword vs gun scene.
A hundred years ago computers were 'electronic brains'.
It goes dormant after loosing its edge and then re-emerges when some research project gains traction again.
logic programming is not directly linked to SQL, and has its own AI term now: https://en.wikipedia.org/wiki/GOFAI.
Humans are intelligent + humans play go => playing go is intelligent
Humans are intelligent + humans do algebra => doing algebra is intelligent
Meanwhile, humans in general are pretty terrible at exact, instantaneous arithmetic. But we aren’t claiming that computers are intelligent because they’re great at it.
Building a machine that does a narrowly defined task better than a human is an achievement, but it’s not intelligence.
Although, in the case of LLMs, in context learning is the closest thing I’ve seen to breaking free from the single-purpose nature of traditional ML/AI systems. It’s been interesting to watch for the past couple years because I still don’t think they’re “intelligent”, but it’s not just because they’re one trick ponies anymore. (So maybe the goalposts really are shifting?) I can’t quite articulate yet what I think is missing from current AI to bridge the gap.
"The question of whether a computer can think is no more interesting than the question of whether a submarine can swim." - Edsger Dijkstra
it is really breaking free? so far LLMs in action seem to have a fairly limited scope -- there are a variety of purposes to which they can be applied but it's all essentially the same underlying task
Indeed, if AI was an algorithm, imagine what would it feel like to be like one: at every step of your thinking process you are dragged by the iron hand of the algorithm, you have no agency in decision making, for every step is pre-determined already, and you're left the role of an observer. The algorithm leaves no room for intelligence.
But that is an uncomfortable idea for most people.
I agree with you that people don't consider intelligence as fundamentally algorithmic. But I think the appeal of algorithmic intelligence comes from the fact that a lot of intelligent behaviours (strategic thinking, decomposing a problem into subproblems, planning) are (or at least feel) algorithmic.
our brain is mostly scatter-gather with fuzzy pattern matching that loops back on itself. which is a nice loop, inputs feeding in, found patterns producing outputs and then it echoes back for some learning.
but of course most of it is noise, filtered out, most of the output is also just routine, most of the learning happens early when there's a big difference between the "echo" and the following inputs.
it's a huge self-referential state-machine. of course running it feels normal, because we have an internal model of ourselves, we ran it too, and if things are going as usual, it's giving the usual output. (and when the "baseline" is out of whack then even we have the psychopathologies.)
There are really very few situations where people really want voice recognition. The main one is hands free controls when driving or for as a remote control for TV or music.
But it may be a while.
I don't doubt that there were marketing people wanting to attach an AI label to them, but that's just marketing BS.
You can thank social media for dumbing down a human technological milestone in artificial intelligence. I bet if there was social media around when we landed on the moon you’d get a lot of self important people rolling their eyes at the whole thing too.
AI = machine learning
Yes, that makes a convolutional net trained on recognizing digits an AI.
I'm not sure, but I think statistical regression already specifies a simple model (e.g. a linear one), while the approach with neural networks doesn't make such strong assumptions about the model.
I actually think linear regression isn't usually called ML in practice. Anyway, I meant something like neural networks.
The ones that win will win not just on technology, but on talent retention, business relationships/partnerships, deep funding, marketing, etc. The whole package really. Losing is easy, miss out on one of these for a short period of time and you've easily lost.
There is no major moat, except great execution across all dimensions.
"There is, however, one enormous difference that I didn’t think about: You can’t build a cloud vendor overnight. Azure doesn’t have to worry about a few executives leaving and building a worldwide network of data centers in 18 months."
This isn't true at all. There are like 8 of these companies stood up in the last three or four years fueled by massive investment of sovereign funds - mostly the saudi, dubai, northern europe, etc. oil-derived funds - all spending billions of dollars doing exactly that and getting something done.
The real problem is the ROI on AI spending is.. pretty much zero. The commonly asserted use cases are the following:
Chatbots Developer tools RAG/search
Not a one of these is going to generate $10 of additional revenue per sollar spent, nor likely even $2. Optimizing your customer services representatives from 8 conversations at once to an average of 12 or 16 is going to save you a whopping $2 per hour per CSR. It just isn't huge money. And RAG has many, many issues with document permissions that make the current approaches bad for enterprises - where the money is - who as a group haven't spent much of anything to even make basic search work.
I agree with you that ROI on _most_ AI spending is indeed poor, but AI is more than LLM's. Alas, what used to be called AI before the onset of the LLM era is not deemed sexy today, even though it can still make very good ROI when it is the appropriate tool for solving a problem.
Amazon doesn't need to worry about suddenly losing its entire customer base to Alibaba, Yandex, or Oracle.
Companies in user acquisition/growth mode tend to have low internal ROI, but remember both Facebook and Google has the same issue -- then they introduced ads and all was well with their finances. Similar things will happen here.
Basically, I'd argue that LLMs look less like a Web 2.0 social media opportunity and more like Hashicorp or Docker, except with operational expenses running many orders of magnitude higher with costs scaling linearly to revenue.
Why can't these providers access all documents and when answers are prompted, self-censor if the reply has references to documents that the end users do not have access permissions? In fact, I'm pretty sure that's how existing RAGaaS providers are handling document/file permissions.
We are? What innovation?
What do we need innovation for? What present societal problems can tech innovation possibly address? Surely none of the big ones, right? So then is it fit to call technological change - 'innovation'?
I'd agree that LLMs improve upon having to read Wikipedia for topics I'm interested in but would investing billions in Wikipedia and organizing human knowledge have produced a better outcome than relying on a magic LLM? Almost certainly, in my mind.
You see, people are pouring billions into LLMs and not Wikipedia not because it is a better product - but because they foresee a possibility of an abusive monopoly and that really excites them.
That's not innovation - that's more of the same anti-social behaviour that makes any meaningful innovation extremely difficult.
An ad is just a piece of data (eg Bob is selling shovels for $10 + shipping), with the additional metadata that someone really wants you to see it (eg Bob paid Google $100 to tell every person who searched for "shovels" that he's selling shovels for $10 + shipping).
Selling ads is organizing human knowledge.
Yes, you could argue using a market system is not the best way to do this, but it is a way to do it.
At least with the current big AI players there is the potential for differentiation through competition.
Unless there is some similar initiative with the Wikipedias, the problem of single supplier dominance is a difficult one to see as the way forward.
Politics, history, religion and other topics of conversation that are matters of opinion, taste and state sponsored propaganda need to be off limits.
Its mission ought to be to provide a PhD level education in all technical fields, not engage in shortening historical events and/or opinions/preferences/beliefs down to a few pages and disputing which pages need to be left in or out. Let fools engage in that task on their own time.
In gold rush, shell shovels.
In computing power rush, sell energy.
They've gotten better, more efficient, loaded with tech, but are still roughly 4 seats, 4 doors, 4 wheels, driven by petroleum.
I know that this is a massive oversimplification, but I think we have seen the "shape" of LLMs\Gen AI\AI products already and it's all incremental improvements from here on out with more specialization.
We are going to have SUVs, sports cars, and single seater cars, not flying cars. AI will be made more fit for purpose for more people to use, but isn't going to replace people outright in their jobs.
"We've pretty much seen their shape. The IBM PC isn't fundamentally very different from the Apple II. Probably it's just all incremental improvements from here on out."
Putting aside that I fundamentally don't think AGI is in the tech tree of LLM, if you will, that there's no route from the latter to the former: even if there is, even if it takes, I dunno, ten years: I just don't think ChatGPT is a compelling enough product to fund about $70 billion in research costs. And sure, they aren't having to yet thanks to generous input from various commercial and private interests but like... if this is going to be a stable product at some point, analogous to something like AWS, doesn't it have to... actually make some money?
Like sure, I use ChatGPT now. I use the free version on their website and I have some fun with AI dungeon and occasionally use generative fill in Photoshop. I paid for AI dungeon (for awhile, until I realized their free models actually work better for how I like to play) but am now on the free version. I don't pay for ChatGPT's advanced models, because nothing I've seen in the trial makes it more compelling an offering than the free version. Adobe Firefly came to me free as an addon to my creative cloud subscription, but like, if Adobe increased the price, I'm not going to pay for it. I use it because they effectively gave it to me for free with my existing purchase. And I've played with Copilot a bit too, but honestly found it more annoying than useful and I'm certainly not paying for that either.
And I realize I am not everyone and obviously there are people out there paying for it (I know a few in fact!) but is there enough of those people ready to swipe cards for... fancy autocomplete? Text generation? Like... this stuff is neat. And that's about where I put it for myself: "it's neat." OpenAI supposedly has 3.9 million subscribers right now, and if those people had to foot that 7 billion annual spend to continue development, that's about $150 a month. This product has to get a LOT, LOT better before I personally am ready to drop a tenth of that, let alone that much.
And I realize this is all back-of-napkin math here but still: the expenses of these AI companies seem so completely out of step with anything approaching an actual paying user base, so hilariously outstripping even the investment they're getting from other established tech companies, that it makes me wonder how this is ever, ever going to make so much as a dime for all these investors.
In contrast, I never had a similar question about cars, or AWS. The pitch of AWS makes perfect sense: you get a server to use on the internet for whatever purpose, and you don't have to build the thing, you don't need to handle HVAC or space, you don't need a last-mile internet connection to maintain, and if you need more compute or storage or whatever, you move a slider instead of having to pop a case open and install a new hard drive. That's absolutely a win and people will pay for it. Who's paying for AI and why?
The fact that human intelligence exists means that the idea of human level intelligence is not a pipe dream.
The question is whether or not the basic underlying technology of the LLM can achieve that level of intelligence.
In the end the best funded company, Uber, is now the most valuable (~$150B). Lyft, the second best funded, is 30x smaller. Are there any other serious ride sharing companies left? None I know of, at least in the US (international scene could be different).
I don't know how the AI rush will work out, but I'd bet there will be some winners and that the best capitalized will have a strong advantage. Big difference this time is that established tech giants are in the race, so I don't know if there will be a startup or Google at the top of the heap.
I also think that there could be more opportunities for differentiation in this market. Internet models will only get you so far and proprietary data will become more important potentially leading to knowledge/capability specialization by provider. We already see some differentiation based on coding, math, creativity, context length, tool use, etc.
It's a fundamentally different beast from AI companies.
Amazon is the poster-child of that mentality. It spent more than it earned into growth for more than 20 years, got a monopoly on retail, and still isn't the most profitable retail company around.
But I am doubtful that the larger enterprise that is Uber (including all the drivers and their expenses and vehicle depreciation, etc) was profitable. I haven't seen that analysis.
[1] https://www.theverge.com/2024/2/8/24065999/uber-earnings-pro...
AI offers WAAAAYYYYY more money in the future.
Models don’t just compete on capability. Over the last year we’ve seen models and vendors differentiate along a number of lines in addition to capability:
- Safety
- UX
- Multi-modality
- Reliability
- Embeddability
And much more. Customers care about capability, but that’s like saying car owners care about horsepower — it’s a part of the choice but not the only piece.
The UX differences among the models are indeed becoming clearer and more important. Claude’s Artifacts and Projects are really handy as is ChatGPT’s Advanced Voice mode. Perplexity is great when I need a summary of recent events. Google isn’t charging for it yet, but NotebookLM is very useful in its own way as well.
When I test the underlying models directly, it’s hard for me to be sure which is better for my purposes. But those add-on features make a clear differentiation between the providers, and I can easily see consumers choosing one or another based on them.
I haven’t been following recent developments in the companies’ APIs, but I imagine that they are trying to differentiate themselves there as well.
According to who?
I love it. More goodies for us
The ones without money will usually lose because they get less opportunity to get in front of eyeballs. Occasionally they manage it anyway, because despite the myth that the VCs love to tell, they aren't really great at finding and promulgating the best tech.
LLMs are capital intensive. They’re a natural fit for financing.
I’m reminded of slime molds solving mazes [1]. In essence, VC allows entrepreneurs to explore the solution space aggressively. Once solutions are found, resources are trimmed.
[1] https://www.mbl.edu/news/how-can-slime-mold-solve-maze-physi...
Except for all the others.
Getting lucky twice is a row is really really lucky. Getting lucky three times in a row is not more likely because they were lucky two times in a row
Tech companies purchased television away from legacy media companies and added (1) unskippable ads, (2) surveillance, (3) censorship and revocation of media you don't physically own, and now they're testing (4) ads while shows are paused.
There's no excuse for getting fooled again.
They destroyed the Taxi industry, I used to be able to just walk out to the taxi rank and get in the first taxi, but not anymore. Now I have to organize it on an app or with a phone call to a robot, then wait for the car to arrive, and finally I have to find the car among all the others that other people called.
Food delivery used to be done by the restaurants own delivery staff, it was fast, reliable and often free if ordering for 2+ people. Now it always costs extra, and there are even more fees if I want the food while it's still hot. Zero care is taken with the delivery, food/drinks are not kept upright and can be a total mess on arrival. Sometimes it's escaped the container and is just in the plastic bag. I have ended up preferring to go pickup food myself over getting it delivered, even when I have a migraine, it's just gone to shit.
I assume you are talking about airports. Guess what, they still exist in many places. And on the other hand, for US, other than a few big cities, the "normal" taxi experience is that you call a number and maybe a taxi shows up in half an hour. With Uber, that becomes 10 minutes or less, with live map updates. Give me that and I'll be happy to forget about Uber.
Getting a taxi was awful before ride-sharing apps. You'd have to walk to a taxi stop, or wait on the side of the road and hope you could hail one. Once the ride-sharing apps came in, suddenly getting a ride became a lot simpler. Our taxi companies are still alive, though they have their own apps now -- something that wouldn't have happened without competition -- and they also work together with the ride-hailing companies as a provider. You could still hail taxis or get them from stops too, though that isn't recommended given that they might try to run the meter by taking a longer route.
For food delivery, before the apps, most places didn't deliver food. Nowadays, more places deliver. Even if a place already had their own delivery drivers, they didn't get rid of them. We get a choice, to use the app or to use the restaurant's own delivery. Usually the app is better for smaller meals since it has a lower minimum order amount, but the restaurant provides faster delivery for bigger orders.
I can now summon a cab from the comfort of the phone I'm holding, and know that they'll accept my credit card. I know the price before I get in and I know the route they should take. I'm not going to get taken for an unnecessary scenic tourist surcharge detour.
people don't like feeling they got cheated, and pre-uber, taxis did that all the time.
In every city I have lived, that is a good thing, despite all the bad things ridesharing startups may have done.
These rideshare and delivery companies are disgusting and terrible.
Leaves you in a worse position than you started with.
Monetizing all of this is frankly...not my problem.
Very few people manage that - indeed I can't think of anyone. Even movie stars get replaced with other movie stars if they try to charge too much. Certainly everyone in the tech industry (including the CEOs, the VCs, the investors etc.) has a viable substitute.
I see 2 paths: - Consumers - the Google way: search and advertise to consumers - Businesses - the AWS way: attrack businesses to use your API and lock them in
The first is fickle. Will OpenAI become the door to the Internet? You'll need people to stop using Google Search and rely on ChatGPT for that to happen. Will become a commodity. Short term you can charge a subscription but long term will most likely become a commondity with advertising.
The second is tangible. My company is plugged directly to the OpenAI API. We build on it. Still very early and not so robust. But getting better and cheaper and faster over time. Active development. No reason to switch to something else as long as OpenAI leads the pack.
There are so many ways, it makes the question seem nonsensical.
Ways to monetize AI so far:
Metered APIs (OpenAI and others)
Subscription products built on it (Copilot, ChatGPT, etc.)
Using it as a feature to give products a competitive edge (Apple Intelligence, Tesla FSD)
Selling the hardware (Nvidia)
What similar parallel can we think of for AI?
Has it become a verb yet? Waiting to peole to replace "I googled how to..." with "I chatgpted how to...".
Whether something has more name recognition isn’t completely related here. But if that’s what you really mean, as you state, “any common mortal about AI and they'll mention ChatGPT - not Claude, Gemini or whatever else. They might not even know OpenAI. But they do know ChatGPT,” then I mostly agree, but as an outsider it doesn’t seem like this is a robust reason to build on top of.
They're issues of socially and politically organizing how humans and resources and economies interact, and there is nothing about the current crop of AI that suggests they can help with that in any particular capacity. If anything the genAI hunger for energy and chips is kinda looking like a net negative in terms of the first problem.
I don't even really get why the hypemen are pushing these huge global issues as motivators - surely that insane standard of Unrealised Value only highlights how limited the abilities of this technology actually are?
Sam Altman is either seriously drunk on his own success, or he's a snake-oil salesman: https://ia.samaltman.com/ (Could be both; they are not mutually exclusive)
ChatGPT is most popular and often in news, because it is the first of its kind(after siri/cortana) which was accessible to general public. I think, I have seen at least 100+ indie and commercial wrappers/apps offering access to better chat interface and common place to access various models from OAI/Anthropic/Mistral. At least one business I read about this weekend uses OAI API to ingest and summarize extensive mountain of documents to simplify some sort of regulatory certification process for medical device manufacturers. There was another one, which allowed a virtual assistant with basic activity ability(book appointments, write emails, transcribe phone calls, summarize letters etc.)
Things will come eventually, as more and more people get access to the APIs and access becomes cheap(current model of per-token price is still immensely expensive imho) and general people do not understand what context is, what embedding is, what exactly is a model, what is an API.
AI was/is being used massively in medical sector, defense/military, fraud detection(finance) and various other sectors for decades or more. Unfortunately, those applications are behind closed door and do not need to drum up public interest or investments directly as those are already well financed and usually referred to as ML/DL/CV etc.
For example, I was doing an internship many many years back where my employer was using "AI"(though it was not referred to as such) to combine vision, noise, vibration, pressure and hundreds of other data to detect hit-n-run, break-in crimes, structural damage, prediction of collapse of large structures etc. and these were prominently used by insurance as well as some public service offices. These were specialized things, and outside of their closed doors, no one was interested about these.
But here are some examples of things that used to fall under the AI umbrella but don't really anymore:
- Fulltext search with decent semantic hit ranking (Google)
- Fulltext search with word sense disambiguation (Google)
- Fulltext search with decent synonym hits (Google)
- Machine translation
- Text to speech
- Speech to text
- Automated biometric identification (Like for unlocking your phone)
If you're more specifically asking for everyday applications of GPT-style generative large language models, I don't think that's going to happen for cost reasons. These things are still far too expensive for use in everyday consumer products. There's ChatGPT, but it's kind of an open secret that OpenAI is hemorrhaging money on ChatGPT.Only vcs need something new to hype.
https://theonion.com/recession-plagued-nation-demands-new-bu...
Regarding point 6, Amazon invested heavily in building data centers across the US to enhance customer service and maintain a competitive edge. it was risky.
This strategic move resulted in a significant surplus of computing power, which Amazon successfully monetized. In fact, it became the company's largest profit generator.
After all, startups and businesses is all about taking risk, ain't it?
The move of the retail marketplace hosting from seattle to virginia happened alongside, and continued long after, the start of AWS.
It is an utter myth, not promoted by Amazon, that AWS was some scheme to use “surplus computing power” from retail hosting/operations. It was an intentional business case to get in to a different B2B market as a service provider.
Our system works, is AI, is profitable, doing vision. Vision scales. There's a little bit of LLM classification. And robotics also, but this part is not really AI, just a generic industry robot.
AI companies are promising AGI to investors to survive a few more years before they probably collapse and don't deliver on that promise.
LLMs are now a commodity. It's time for startups to build meaningful products with it!
A clever AI algorithm run on rented compute is not a moat.
This is the Red Queen hypothesis in evolution. You have to keep running faster just to stay in place.
On it's face, this does seem like a sound argument that all the $$ following LLMs is irrational:
1. No matter how many billions you pour into your model, you're only ever, say, six months away from a competitor building a model that's just about as good. And so you already know you're going to need to spend an increased number of billions next year.
2. Like the gambler who tries to beat the house by doubling his bet each time, at some point there must be a number where that many billions is considered irrational by everybody.
3. Therefore it seems irrational to start putting in even the fewer billions of dollars now, knowing the above two points.
This doesn’t follow. One, there are cash flows you can extract in the interim—standing in place is potentially lucrative.
And two, we don’t know if the curve continues until infinity or asymptotes. If it asymptotes, being the first at the asymptote means owning a profitable market. If it doesn’t, you’re going to get AGI.
Side note: bought but haven’t yet read The Red Queen. Believe it was a comment of yours that lead me to it.
For a market to be both winner-take-all as well as lucrative, you need some kind of a feedback cycle from network effets, economies of scale, or maybe even really brutal lock-in. For example, for operating systems applications provide a network effect -- more apps means an OS will do better, a higher install base for an OS means it'll attract more apps. For social networks it's users.
One could imagine this being the case for LLMs; e.g. that LLMs with a lot of users can improve faster. But if the asymptote has been reached, then by definition there's no more improvement to be had, and none of this matters.
Larger user base = increased feedback = improved quality of answers = moat
1. Training the next generation of models
2. Providing worldwide scalable infrastructure to serve those models (ideally at a profit)
It's hard enough to accomplish #1, without worrying about competing against the hyperscalers on #2. I think we'll see large licensing deals (similar to Anthropic + AWS, OpenAI + Azure) as one of the primary income sources for the model providers.
With the second (and higher margin) being user facing subscriptions. Right now 70% of OpenAI's revenue comes from chatgpt + enterprise gpt. I imagine Anthropic is similar, given the amount of investment in their generative UI. At the end of the day, model providers might just be consumer companies.
Safely and effectively, that is. Dangerous and inappropriate is obviously a much wider set of possibilities.
The cardiologist checks the ECG, compare with the LLM results and checks the difference. If it can reduce error rate by like 10%, that's already really good.
My current stance on LLM is that it's good for stuff which is painful to generate, but easy to check (for you). It's easier/faster to read an email than to write it. If you're a domain expert, you can check the output, and so on. The danger is in using it for stuff you cannot easily check, or trusting it implicitly because it is usually working.
I guess their secret sauce was just so good and so secret that neither established players nor copycat startups were able to replicate it, the way it happened with ChatGPT? Why is the same not the case here, is it just because the whole LLM thing grew out of a relatively open research culture where the fundamentals are widely known? OTOH PageRank was also published before the founding of Google.
I'd be curious to hear if anyone has theories or insight here.
Conversely, the 2020s has execs and other rich individuals doomscrolling LinkedIn and thinking that if they don't invest in the latest crypto/quantum/genai crap, they're missing out on a vital element of retaining competitive edge.
Microsoft's been trying to ram essentially that down my throat for the better part of a year now, and it's mostly convinced me that the answer is "no". I don't want to have arbitrary conversations with my computer.
I still just want the same thing I've been wanting from my digital assistant for 30 years now: fewer "eat up Martha" moments, and handling more intents so that I can ask "When does the next east-bound bus come?" and it stops answering questions like "Will it rain today?" as if I had asked "Is it raining right now?". None of those are particularly appropriate problems for a GPT-style model.
Set two timers, one for 20 minutes, and another for 50 minutes.
or
Turn off the lights in my living room, and my office.
That's about as advanced as I want my home assistants to be though.
I asked Google Gemini:
Tailored Ads: Create highly personalized ads based on this data, increasing the likelihood of clicks and conversions.
Could probably also have a version that goes something like
> you asked about $SUBJECT. Here is some advice on the subject, and a few links to helpful products and services.
I see no reason why AI will be particularly different. It seems difficult to make the case AI is useless, but it’s also not particularly mature with respect to fundamental models, tool chains, business models, even infrastructure.
In both cases speculative capital flowed into the entire industry, which brought us losers like pets.com but winners like Amazon.com, Netflix.com, Google.com, etc. Which of the AI companies today are the next generation of winners and losers? Who knows. And when the music stops will there be a massive reckoning? I hope not, but it’s always possible. It probably depends on how fast we converge to “what works,” how many grifters there are, how sophisticated equity investors are (and they are much more sophisticated now than they were in 1997), etc.
1. It's not that easy to switch between providers. There's no lock-in of course, but once you build a bunch of code that is provider specific (structured outputs, prompt caching, json mode, function calls, prompts designed for a specific provider, specific tools used by the openai Assistant, etc) then you need a good reason to switch (like a much better or cheaper model)
2. All of these companies do try to build some echo system around them, esp in the enterprise. The problem is that Google and Microsoft have a huge advantage here cause they have all the integrations
3. The consumer side. It's not just LLM. It's image, video, voice, and many more. You cannot ignore that ChatGPT can rival Google in a few years in terms of usage. As long as they can deliver good models, users are not going to switch so quickly. It's a huge market, just like Google. Pretty much, everyone in the world is going to use ChatGPT or some alternative in the next few years. My 9 year old and her friends already use it. No reason why they cannot monetize their huge user base like Google did.
And at some point one of these companies will reach the point it does not need as many employees. And has a model capable of efficiently incorporating new learning without having to reset and relearn from scratch.
That is what AGI is.
Computing resources for inference and incremental learning will still be needed, but when the AGI itself is managing all/much of that, including continuing to find efficiencies, ... profitably might be unprecedented.
The speed of advance over the last two decades has been steady and exponential. There are not many (or any) credible signals that a technical wall is about to be encountered.
Which is why I believe that I, I by myself, might get there. Sort of, kind of, probably not, probably just kidding. Myself.
--
Another reason companies are spending billions is to defend their existing valuations. Google's value could go to zero if they don't keep up. Other companies likewise.
It is the new high stakes ante for large informational/social service relevance.
My guess is that profitability is becoming increasingly difficult and nobody knows how yet… or whether it will be possible.
Seems like the concentration of capital is forcing the tech industry to take wilder and more ludicrous bets each year.
Missed one I think... the expertise accumulated in building the prior generation models, that are not themselves that useful anymore.
Yes, it's true that will be lost if everybody leaves, a point he briefly mentions in the article. But presumably AWS would also be in trouble, sooner or later, if everybody who knows how things work left. Retaining at least some good employees is tablestakes for any successful company long-term.
Brand and inertia also don't quite capture the customer lock-in that happens with these models. It's not just that you have to rewrite the code to interface with a competitor's LLM; it's that that LLM might now behave very differently than the one you were using earlier, and give you unexpected (and undesirable) results.
I know we aren't in the 90s. I know that the cost of successive process nodes has grown exponentially, even when normalizing for inflation. But, still. I'd be wary of betting the farm on AI being eternally confined to giant, expensive special-purpose hardware.
This stuff is going to get crammed into a little special purpose chip dangling off your phone's CPU. Either that, or GPU compute will become so commodified that it'll be a cheap throw-in for any given VPS.
The idea that you need huge amounts of compute to innovate in a world of model merging and activation engineering shows a failure of imagination, not a failure to have the necessary resources.
PyReft, Golden Gate Claude (Steering/Control Vectors), Orthogonalization/Abliteration, and the hundreds of thousands of Lora and other adapters available on websites like civit.ai is proof that the author doesn't know what they're talking about re: point #2.
And I'm not even talking about the massive software/hardware improvements we are seeing for training/inference performance. I don't even need that, I just need evidence that we can massively improve off the shelf models with almost no compute resources, which I have.
That's a fundamental misunderstanding of the (especially) US startup culture in the last maybe 10-20 years. Only very rarely is the goal of the founders and angel investors to build an actual sustainable business.
In most cases the goal is to build enough perceived value by wild growth financed by VC money & by fueling hype that an subsequent IPO will let the founders and initial investors recoup their investment + get some profit on top. Or, find someone to acquire the company before it reaches the end of its financial runway.
And then let the poor schmucks who bought the business hold the bag (and foot the bill). Nobody cares if the company becomes irrelevant or even goes under at that point anymore - everyone who did has has recouped their expense already. If the company stays afloat - great, that's a bonus but not required.
1. it takes huge and increasing costs to build newer models, models approached asymptote 2. a startup can take an open-source model and get you out of business in 18 months (with CocaCola example)
The size of LLM is what protects them from being attacked by startups. Microsoft's operating profit in 2022 was $72B, which is 10x bigger than the running cost of OpenAI. And if 2022 was too successful, profits of $44B still dwarf OpenAI.
If OpenAI manages to ramp up investment like Uber, it may stay alive, otherwise it's tech giants that can afford running some LLM. ...if people will be willing to pay for this level of quality (well, if you integrate it into MS Word, they actually may want it).
Do you think it is plausible that they will be able to get enough people to pay for OpenAI that Netflix? They would need much more revenue than Netflix to make it viable. Considering that there are other options available, like Google’s, MSFT’s, Anthropic’s, etc.
As a business model, this is all very suspect. GPT5 has had multiple delays and challenges. What if “marginally better” is all we are going to get?
Just food for thought.
(1) High integration (read: switching) costs: any deployment of real value is carefully tested and tuned for the use-case (support for product x, etc.). The use cases typically don't evolve that much, so there's little benefit to re-incurring the cost for new models. Hence, customers stay on old technology. This is the rule rather than the exception e.g., in medical software.
(2) The Instagram model: it was valuable with a tiny number of people because they built technology to do one thing wanted by a slice of the market that was very interesting to the big players. The potential of the market set the time value of the delay in trying to replicate their technology, at some risk of being a laggard to a new/expanding segment. The technology gave them a momentary head start when it mattered most.
Both cases point to good product-market fit based on transaction cost economics, which leads me to the "YC hypothesis":
The AI infrastructure company that best identifies and helps the AI integration companies with good product-market fit will be the enduring leader.
If an AI company's developer support consist of API credits and online tutorials about REST API's, it's a no-go. Instead, like YC and VC's, it should have a partner model: partners use considerable domain skills to build relationships with companies to help them succeed, and partners are selected and supported in accordance with the results of their portfolio.
The partner model is also great for attracting and keeping the best emerging talent. Instead of years of labor per startup or elbowing your way through bureaucracies, who wouldn't prefer to advise a cohort of the best prospects and share their successes? Unlike startup's or FAANG, you're rewarded not for execution or loyalty, but for intelligence in matching market needs.
So the question is not whether the economics of broadcast large models work, but who will gain the enduring advantage in supporting AI eating the software that eats the world?
Or search/social media where people are happy to pay $0 to use it in exchange for ads.
Sure some people are paying now, but its nowhere near the cost of operating these models, let alone developing them.
Also the economics may not accrue to the parts of the stack people think. What if the model is commodity and the real benefits accrue to the GOOG/AAPL/MSFT of the world that integrate models, or to the orgs that gatekept their proprietary data properly and now can charge for querying it?
This may become increasingly the case as models get smarter, but it’s often not the case right now. It’s more likely to be a few lines of code, a bunch of testing, and then a bunch of prompt tweaking iterations. Even within a vendor and model name, it’s a good idea to lock to a specific version so that you don’t wake up to a bunch of surprise breakages when the next version has different quirks.
that's like a better half of the entire Apple business model that brought them success, find the next hot thing (like a capacitive touchscreen or high density displays), make exclusive deals with hardware providers so you can release a device using that new tech and feed on it for some time till the the upstream starts leaking tech left and right and you competition can finally catch up
My prediction is that the thin layer built on top of LLMs will be eaten up starting from the most profitable products.
By using inference APIs you are doing their market research for free.
The question is how can you use in-context learning to optimize the model weights. It’s a fun math problem and it certainly won’t take a billion dollar super computer to solve it.
https://arxiv.org/abs/2407.10930 https://arxiv.org/abs/2006.04439
What about: Lobby for AI regulations that prevent new competitors from arising, and hopefully kills of a few?
it's a telling quote he chose: if you think AI is over invested, you should short AI companies, and that's where the quote comes from, problem is, even if you're right the market can stay irrational longer than you can afford to hold your short.
I don’t think most people looking to build an AI company want to build an LLM and call it a company.
The upside risk is premised on this point. It'll get so cost prohibitive to build frontier models that only 2-3 players will be left standing (to monetize).
Not that physical/financial constraints are unimportant, but they often can be mitigated in other ways.
Some background: I was previously at one of these companies that got hoovered up in the past couple years by the bigs. My job was sort of squishy, but it could be summarized as 'brand manager' insofar as it was my job to aide in shaping the actual tone, behaviors, and personality of our particular product.
I tell you this because in full disclosure, I see the world through the product/marketing lens as opposed to the engineer lens.
They did not get it.
And by they I mean founds whose names you've heard of, people with absolute LOADS of experience in building a shipping technology products. There were no technical or budgetary constraints at this early stage, we were moving fast and trying shit. But they simply could not understand why we needed to differentiate and how that'd make us more competitive.
I imagine many technology companies go through this, and I don't blame technical founders who are paranoid about this stuff; it sounds like 'management bullshit' and a lot of it is, but at some point all organizations who break even or take on investors are going to be answerable to the market, and that means leaving no stone unturned in acquiring users and new revenue streams.
All of that to say, I do think a lot of these AI companies have yet to realize that there's a lot to be done user experience-wise. The interface alone - a text prompt(!?) is crazy out-of-touch to me. The fact that average users have no idea how to set up a good prompt and how hard everyone is making it for them to learn about that.
All of these decisions are pretty clearly made by someone who is technology-oriented, not user-oriented. There's no work I'm aware of being done on tone, or personality frameworks, or linguistics, or characterization.
Is the LLM high on numeracy? Is it doing code switching/matching, and should it? How is it qualifying its answers by way of accuracy in a way that aids the user learning how to prompt for improved accuracy? What about humor or style?
It just completely flew over everyone's heads. This may have been my fault. But I do think that the constraints you see to growth and durability of these companies will come down to how they're able to build a moat using strategies that don't require $$$$ and that cannot be easily replicated by competition.
Nobody is sitting in the seats at Macworld stomping their feet for Sam Altman. A big part of that is giving customers more than specs or fiddly features.
These companies need to start building a brand fast.
I don't see it this way. In plumbing I could have chosen to use 4" pipe throughout my house. I chose 3". Heck, I could have purchased commercial pipe that's 12", or even 36". It would have changed a lot of the design of my foundation.
Just because there is something much bigger and can handle a lot more poop, doesn't mean it's going to be useful for everyone.
LLMs as coding assistants seem to be great. Let’s say that every working programmer will need an account and will pay $10/month (or their employer will).. what’s a fair comp for valuation? GitHub? That’s about $10Bn. Atlassian? $50Bn
The “everything else” bin is hard to pin down. There are some clear automation opportunities in legal, HR/hiring, customer service, and a few other fields - things that feel like $1-$10Bn opportunities.
Sure, the costs are atrocious, but what’s the revenue story?
Replace coding assistant with artists and you have the vibe of AI 2 years ago.
The issue is that these models are easy to make (if expensive) so the open source community (of which many, maybe most, are programmers themselves) will likely eat up any performance moat given enough time.
This story already played out with AI art. Nothing beats SD and comfyUI if you really need high quality and control.
Here's a counterargument.
> In other words, the billions that AWS spent on building data centers is a lasting defense. The billions that OpenAI spent on building prior versions of GPT is not, because better versions of it are already available for free on Github.
The money that OpenAI spends on renting GPUs to build the next model is not what builds the moat. The moat comes from the money/energy/expertise that OpenAI spends on the research and software development. Their main asset is not the current best model GPT-4; it is the evolving codebase that will be able to churn out GPT-5 and GPT-6. This is easy to miss because the platform can only churn out each model when combined with billions of dollars of GPU spend, but focusing on the GPU spend misses the point.
We're no longer talking about a thousand line PyTorch file with a global variable NUM_GPUs that makes everything better. OpenAI and competitors are constantly discovering and integrating improvements across the stack.
The right comparison is not OpenAI vs. AWS, it's OpenAI vs. Google. Google's search moat is not its compute cluster where it stores its index of the web. Its moat is the software system that incorporates tens of thousands of small improvements over the last 20 years. And similar to search, if an LLM is 15% better than the competitors, it has a good shot at capturing 80%+ of the market. (I don't have any interest in messing around with a less capable model if a clearly better one exists.)
Google was in some sense "lucky" that when they were beginning to pioneer search algorithms, the hardware (compute cluster) itself was not a solved problem the way it is today with AWS. So they had a multidimensional moat from the get-go, which probably slowed early competition until they had built up years' worth of process complexity to deter new entrants.
Whereas LLM competition is currently extremely fierce for a few reasons: NLP was a ripe academic field with a history of publishing and open source, VC funding environment is very favorable, and cloud compute is a mature product offering. Which explains why there is currently a proliferation of relatively similar LLM systems:
> Every LLM vendor is eighteen months from dead.
But the ramp-up time for competitors is only short right now because the whole business model (pretrain massive transformers -> RLHF -> chatbot interface) was only discovered 18 months ago (ChatGPT launched at the end of 2022) - and at that point all of the research ideas were published. By definition, the length of a process complexity moat can't exceed how long the incumbent has been in business! In five years, it won't be possible to raise a billion dollars and create a state of the art LLM system, because OpenAI and Anthropic will have been iterating on their systems continuously. Defections of senior researchers will hurt, and can speed up competitor ramp-time slightly, but over time a higher proportion of accumulated insights is stored in the software system rather than the minds of individual researchers.
Let me emphasize: the billions of dollars of GPU spend is a distraction; we focus on it because it is tangible and quantifiable, and it can feel good to be dismissive and say "they're only winning because they have tons of money to simply scale up models." That is a very partial view. There is a tremendous amount of incremental research going on - no longer published in academic journals - that has the potential to form a process complexity moat in a large and relatively winner-take-all market.
The business model for those is to produce amazing LLM models that are hard to re-create unless you have similar resources and then make money providing access, licensing, etc.
What are those resources? Compute, data, and time. And money. You can compensate for lack of time by throwing compute at the problem or less/more data. Which is a different way of saying: spend more money. So, it's no surprise that this space is dominated by trillion dollar companies with near infinite budgets and a small set of silicon valley VC backed companies that are getting multi billion dollar investments.
So the real question is whether these companies have enough of a moat to defend their multi billion dollar investments. The answer seems to be no. For three reasons: hardware keeps getting cheaper, software keeps getting better, and using the models is a lot cheaper than creating them.
Creating GPT-3 was astronomically expensive a few years ago and now it is a lot cheaper by a few orders of magnitude. GPT-3 is of course obsolete now. But I'm running Llama 3.2 on my laptop and it's not that bad in comparison. That only took 2 years.
Large scale language model creation is becoming a race to the bottom. The software is mostly open source and shared by the community. There is a lot of experimentation happening but mostly the successful algorithms, strategies, and designs are quickly copied by others. To the point where most of these companies don't even try to keep this a secret anymore.
So that means new, really expensive LLMs have a short shelf life where competitors struggle to replicate the success and then the hardware gets cheaper and others run better algorithms against whatever data they have. Combine that with freely distributed models and the ability to run them on cheap infrastructure and you end up with a moat that isn't that hard to cross.
IMHO all of the value is in what people do with these models. Not necessarily in the models. They are enablers. Very expensive ones. Perhaps a good analogy is the value of Intel vs. that of Microsoft. Microsoft made software that ran on Intel chips. Intel just made the chips. And then other chip manufacturers came along. Chips are a commodity now. Intel is worth a lot less than MS. And MS is but a tiny portion of the software economy. All the value is in software. And a lot of that software is OSS. Even MS uses Linux now.
The recent advances are truly jaw dropping. It absolutely merits being investigated to the hilt. There is a very good chance that it will end up being a net profit.
But intuitively they don't feel to me like they're getting more human. If anything I feel like the recent round of "get it to reason aloud" is the opposite of what makes "general intelligence" a thing. The vast majority of human behavior isn't reasoned, aloud or otherwise.
It'll be super cool if I'm wrong and we're just one algorithm or extra data set or scale factor of CPUs away from It, whatever It turns out to be. My intuition here isn't worth much. But I wouldn't be surprised if it was right despite that.