I wonder how Paul Graham thinks of Sam Altman basically copying Cursor and potentially every upstream AI company out of YC, maybe as soon as they launch on demo day.
Is it a retribution arc?
I wonder how Paul Graham thinks of Sam Altman basically copying Cursor and potentially every upstream AI company out of YC, maybe as soon as they launch on demo day.
Is it a retribution arc?
If OpenAI can copy Cursor, so can everyone else.
Good prompts may actually have a moat - a complex agent system is basically just a lot of prompts and infra to co-ordinate the outputs/inputs.
The second part of that statement (is wrong and) negates the first.
Prompts aren’t a science. There’s no rationale behind them.
They’re tricks and quirks that people find in current models to increase some success metric those people came up with.
They may not work from one model to the next. They don’t vary that much from one another. They, in all honesty, are not at all difficult or require any real skill to make. (I’ve worked at 2 AI startups and have seen the Apple prompts, aider prompts, and continue prompts) Just trial and error and an understanding of the English language.
Moreover, a complex agent system is much more than prompts (the last AI startup and the current one I work at are both complex agent systems). Machinery needs to be built, deployed, and maintained for agents to work. That may be a set of services for handling all the different messaging channels or it may be a single simple server that daisy chains prompts.
Those systems are a moat as much as any software is.
Prompts are not.
If one spends a lot of time building an application to achieve an actual goal they'll realize the prompts make a gigantic difference and it takes an enormous amount of fiddly, annoying work to improve. I do this (and I built an agent system, which was more straightforward to do...) in financial markets. It so much so that people build systems just to be able to iterate on prompts (https://www.promptlayer.com/).
I may be wrong - but I'll speculate you work on infra and have never had to build a (real) application that is trying to achieve a business outcome. I expect if you did, you'd know how much (non sexy) work is involved on prompting that is hard to replicate.
Hell, papers get published that are just about prompting!
https://arxiv.org/abs/2201.11903
This line of thought effectively led to Gpt-4-o1. Good prompts -> good output -> good training data -> good model.
Important and easy to make are not the same
I never said prompts didn’t matter, just that they’re so easy to make and so similar to others that they aren’t a moat.
> I may be wrong - but I'll speculate you work on infra and have never had to build a (real) application that is trying to achieve a business outcome.
You’re very wrong. Don’t make assumptions like this. I’ve been a full stack (mostly backend) dev for about 15 years and started working with natural language processing back in 2017 around when word2vec was first published.
Prompts are not difficult, they are time consuming. It’s all trial and error. Data entry is also time consuming, but isn’t difficult and doesn’t provide any moat.
> that is hard to replicate.
Because there are so many factors at play _besides prompting. Prompting is the easiest thing to do in any agent or RAG pipeline. it’s all the other settings and infra that are difficult to tune to replicate a given result. (Good chunking of documents, ensuring only high quality data gets into the system in the first place, etc)
Not to mention needing to know the exact model and seed used.
Nothing on chatgpt is reproducible, for example, simply because they include the timestamp in their system prompt.
> Good prompts -> good output -> good training data -> good model.
This is not correct at all. I’m going to assume you made a mistake since this makes it look like you think that models are trained on their own output, but we know that synthetic datasets make for poor training data. I feel like you should know that.
A good model will give good output. Good output can be directed and refined with good prompting.
It’s not hard to make good prompts, just time consuming.
They provide no moat.
> but we know that synthetic datasets make for poor training data
This is a silly generalization. Just google "synthetic data for training LLMs" and you'll find a bunch of papers on it. Here's a decent survey: https://arxiv.org/pdf/2404.07503
It's very likely o1 used synthetic data to train the model and/or the reward model they used for RLHF. Why do you think they don't output the chains...? They literally tell you - competitive reasons.
Arxiv is free, pick up some papers. Good deep learning texts are free, pick some up.
Training a model on synthetic data (obviously) increases bias present in the initial dataset[1], making for poor training data.
IIRC (this subject is a little fuzzy for me) using synthetic data for RLHF is equivalent to just using dpo, so if they did RLHF it probably wasn’t with synthetic data. They may have gone with dpo, though.
Researchers are using synthetic data to train LLMs, especially for fine tuning, and especially instruct fine tuning. You are not up to date with recent work on LLMs.
Neither was I.
> "synthetic data is bad“
I never said that… I said that it makes for poor training data, which it does.
> Researchers are using synthetic data to train LLMs, especially for fine tuning, and especially instruct fine tuning
Then those researchers are training with subpar datasets as the bias in that data will be compounded.
It’s a trade off since there’s only so much fresh data in form you want. If they could use entirely non synthetic data, I’m sure they would.
And again, you’re choosing to focus on this one point rather than my main point that prompt provide no moat.
> You are not up to date with recent work on LLMs.
There you go again making assumptions…
I think I’m done with this conversation though.
OpenAI are in the money making business. They don’t care about no AGI. They’re experts who know where the limits are at the moment.
We don’t have the tools for AGI any more than we do for time travel.
Your brain is an existential proof that general intelligence isn't impossible.
Figuring out the special sauce that makes a human brain able to learn so much so easily? Sure that's hard, but evolution did it blindly, and we can simulate evolution, so we've definitely got the tools to make AGI, we just don't have the tools to engineer it.
I tried using aider but either my local LLM is too slow or my software projects requires context sizes so large they make aider move at a crawl.
I'd added seperate design/implementation agents before that was added to Aider https://aider.chat/2024/09/26/architect.html
The other different is I have a file selection agent and a code review agent, which often has some good fixes/improvements.
I use both, I'll use Aider if its something I feel it will right the first time or I want control over the files in the context, otherwise I'll use the agent in Sophia.
Met a guy who got brought in by Amazon after they hit 8 figures in sales, wined and dined, then months later Amazon launched competing product and locked them out of their accounts, cost them 9 figures.
You mean downstream.