They're all stealing your IP and selling it back to your competitors in the form of tokens.
They're all stealing your IP and selling it back to your competitors in the form of tokens.
Such a ridiculous stance: "I want LLMs to code for me, but I want them to be trained on other people's code, not mine, duh".
> I have absolutely zero interest in free. I honestly don't think I'm even remotely in the same demographic as people using free tiers / models. I want to pay. I don't want my data used for training...
They want to use LLMs trained on others code but don't want to contribute with their own.
Not casting judgement, just pointing out.
(Disclaimer: Not speaking for or about my current employer, just a general industry observation.)
Who ever said that? Have you actually heard that from your fellow programmers in real life?
If the code I wrote actually made even the slightest discernible difference in LLMs I'd be so honored. But it won't happen, as it's just 0.00001% of all the training data.
> But it won't happen, as it's just 0.00001% of all the training data.
Are you familiar with Tragedy of the Commons?
Cursor users are willfully providing it by using their product. Not unlike uploading a personal photo or video to social media -- that's not yours anymore. You gave all rights away when you put it on their servers.
99% of users are unaware of that, so it’s false to say they "are willfully providing it".
I've granted them a limited license to use it.
> that's not yours anymore.
Not by any definition in the contract or in law is this true.
> You gave all rights away when you put it on their servers.
I gave away some rights. I also got something in return. Attention. And at the end of the day I'm completely entitled to turn around and sell copies of this work for profit. The only thing I can't do is sell an /exclusive/ license because that is no longer available.
None of this provides any implication for people who upload code to their own websites. Which these rapacious LLMs bots happily index, sometimes to the extent they actually crush the site, or create unusual costs for the owner.
Finally none of these LLM companies tell you where the source came from. Whether it is copyrighted, whether ownership rights are retained, or whether the code can be used publicly or not, and if so, which license it's covered by.
You're using the lens of social media contracts to understand something far larger and more important. It's lead you to some bizarre conclusions and huge oversites.
What if someone steals my work and then uploads it to facebook and claims it as their own? Do the rights no longer exist because it got uploaded to Meta?
Personally I like to use a stable IDE. I used Cursor for a couple of days and then went back to VS Code, largely due to Cursor pushing agentic first approach with V3 update.
The users of those tools are stealing too. The model is trained on free software licensed under specific terms and the output of a prompt will strip those licenses and their terms.
As I said, I founded and run a AI lab. Now get over yourself.
Anyway like I said you can just test it out yourself and find out that I'm correct. Every skilled programmer already knows this and can predict what kind of complexity an LLM won't be able to handle. And anyone working on LLMs should know that they are completely dependent on their training data. The entire scaling hypothesis was based on this.
Regardless, your claim is "LLMs being able to one-shot any higher-order complexity is entirely dependent on it already being in the training data." which is currently unfalsifiable, and known by every AI researcher to be most likely wrong. You've also said it's "been demonstrated". So stop wasting people's time and link to the demonstration instead of hand-waving a "do it yourself you're smart".
Good lord, it's like talking to a bloody climate change denier, I swear.
You haven't made a single technical point outside your bloviated claims to authority. I can safely assume you're a fraud because this is exactly how frauds speak. Actual scientists and engineers don't argue from authority they go and test the hypothesis for themselves, the fact that you balk at my suggestion to do this is amusing.
I said you can do a basic test because this is the best way to see it directly for yourself. It's very easy to do especially for an eminent machine learning visionary leader as yourself. You ask the LLM to produce two apps of similar complexity and technical challenge from a software perspective, one that is already in its training data and one that isn't and see which its more successful at. This isn't some controversial take, nor is it "unfalsifiable".
There's also hundreds of benchmarks demonstrating where the limitations are for LLMs, or you can study the progression of LLMs in mathematics and where the gains have been made and see that this also agrees with me. You can watch Chris Hay's videos demonstrating exactly how LLMs perform math layer by layer. Why is everyone using LLMs for search? Because it's an extremely efficient compression of all its training data. Did they figure out the Studio Ghibli art-style all on their own spontaneausly? No, they were trained on Studio Ghibli content. There's so many ways to come to this conclusion. But you seem to be too busy sniffing your own farts to be interested in learning anything about the field though.
I mean, are you seriously trying to back off that original ridiculous claim into a "code in the training data is more likely to appear in the output than code that isn't"? And I'm the fraud? As I said before, get over yourself.
And yes, you are obviously a fraud.