You can't build a moat with AI
generatingconversation.substack.com
generatingconversation.substack.com
It's more than data. Steven Jobs used to tell the founders of the Segway that their secret technologies would leak sooner or later, even if they set up their factories in the middle of a desert in Nevada. His advice to the founders was that they needed to build a product that users couldn't get away from even if all the competitors had the same technology (or so I remember). That could be a process, an ecosystem, market penetration with an amazing supply chain (think about the largest seller of straws in the world. The unit price of a single straw is close to 0, which is really hard to achieve), and etc.
To me, data will just be one key component of the moat but not all. The moat of AI is the same ol' entire ecosystem: an infrastructure that is so efficient that the company can keep driving down the unit price of hosting the AI models; a fabulous culture to enable the company to keep churning out improvements; an extensive data platform and the associated process and sources that keeps provisioning quality data, a number of killer applications...
This sounds encouraging at first glance, but even more demoralizing when you think about it. It doesn’t matter how clever or smart you are over your competitors. If you don’t have the data, you don’t stand a chance. And of course the incumbents have the data not you, the entrepreneur
Whatever thing people build, will be based on data that exists, not data that will be created afterwards. And guess what, you don't have that data.
On the yet to exist data?
Customized small models will outperform larger general models for your specific use case.
And of course, not if you care about token throughput more than fancy abilities. Or price for that matter.
So for many if not most businesses needs GPT-4 isn't the best tool out there, and GPT-5 is the canonical example of a vaporware right now.
While it's possible the gap between GPT5 and 4 is as big as between 4 and 3, it's unlikely. The gap between 2 and 3 was much larger as the one between 3 and 4 (and similarly between 1 and 2).
Also, it's not clear that GPT5 will do this in an *economical way* once the spigot of investor money stops.
But small customized models seem to perform close to as well as large general ones.
This to me is why I am actually bullish on building ‘LLM wrappers’ that are sustainable businesses. While you might not be developing new technology, you can be adding value to a client who can use chat gpt and “talk to your docs” but not the surrounding work to make it frictionless for their workflow or company and actually build a business. It’s entirely possible that AI will get so good that it makes it easy to all the surrounding tasks too.
They have all the data, and no one else even comes close.
Honestly I think they might have more useful data than Google, given Bing knows more or less that same as GoogleBot. Meta doesn't come close, unless you want your LLM to be purely conversational.
Microsoft does seem active in this, e.g. https://microsoft.github.io/presidio/
Moreover, that isn't even the problem here. Suppose your company has a trade secret. You know how to manufacture widgets more efficiently than your competitors. If Microsoft produces a model that will now tell your competitors your secret process that it learned from your internal emails, it's completely irrelevant whether they stripped the PII out of your emails first.
Unlike googles victims (individuals) corporations can and do fight back when someone plays it fast & loose with their confidential coms
This post seems to misunderstand what a "moat" is in a business sense (and unfortunately a lot of AI hypesters on social media do as well). The fact that LLMs are becoming a commodity was the point of the original "OpenAI has no moat" memo by Google, which has proven to be accurate.
And yet, over and over, I see products with output that clearly comes from simple and frankly lazy prompting. You can do a lot with prompting, but engineerings are not putting in the work! (If an engineer is even the right person... probably not, given any specific application of an LLM.)
Prompting also isn't so reductive that you just write an evaluation and then iterate on the prompt until you satisfy the evaluation. Prompting is a co-creative exercise between the LLM, the domain expert, the product, and the user. And sure "data" fits in there, as well as relationships, comprehensibility, workflows, etc etc... the AI component is just a small piece of any full application.
also, it's not "an order of magnitude" more expensive