Turn off telemetry, delete ChatGPT conversations, and I like having an LLM for GitHub
versus
Give us all your personal data and we're sharing it with ad partners, your personal data is our number one product and secondary products are only made to support the first endeavor.
What do you mean by this? What kind of information do advertisers get to receive?
Advertisers don't usually receive detailed information, instead they are promised that their advertisements are targeted correctly. To do that the data about people is used.
But with this data Corporate gets too much power. Corporate needs to follow Human Rights, like states. One aspect is that there is no due process in case of conflicts, like someone being wrongly excluded from using a Corporate service. We also need transparency about how this data is used exactly and who else gets views.
Europe is trying to establish transparency with the GDPR, however is foiled by malicious compliance by Corporate. It does not help that Europe is clumsy.
Also, they sell provide it governments in the form of warrant requests.
For more info, read up in google ad auction fraud, and google location warrant.
Microsoft is actually worse though. For one thing, Windows’ default telemetry level lets them pull files off your machine. Here’s a Microsoft apologist explaining it:
https://www.zdnet.com/article/windows-10-telemetry-secrets/
The US CLOUD Act (which Microsoft lobbied for) means that they have to pull that data in response to US warrant requests, regardless of what country you are in.
Also, they already merged mojang (minecraft) into the Microsoft login umbrella. It is clearly a dry run for when they merge linkedin and github in, so they can provide better worker surveillance capabilities to their corporate customers.
I may be working from old information, but doesn't Windows 10 still have unstoppable, and unauditable, telemetry?
Use a PiHole?
Annoying!
Google and Microsoft are different. The same: I am reluctantly using products from both. The different: Their history and their products.
However I don't think that you should differentiate your trust between them.
We should think about why Corporate does enshittification and how to prevent it.
That LLM was trained on FOSS software that others like myself wrote and did not consent to have used in such a manner. Even if you're also a FOSS contributor and you're fine with it, I'm not.
So you like it, but you're not the party whose consent matters. You're the customer consuming the laundered goods from the fence.
The attribution licenses such as ALv2, BSD, MIT, etc. which are generally considered to be "permissive" still require attribution, and this condition is not upheld by popular LLMs.
Of course, the copied material has to be judged substantial enough for the license to apply in a court of law, which is a human-arbitrated threshold.
It will most likely be a copyleft author who eventually brings a court case, but the attribution licenses do not allow for LLM usage either unless the LLM provides attribution and preserves notices as spelled out in the licenses.
I think it is pretty obvious that lots of GPL authors wouldn’t want their code to be productized in ways that don’t contribute back to the community, because they had the ability to use a BSD or MIT license and didn’t take it.
Legally, of course, LLMs are not people. They don’t have the same rights as people, and it isn’t obvious whether or not they can legally generate new IP. They operate by a complicated but essentially mechanical process, and it is pretty novel to say that such a process could be used to remove copyrights.
But the defendant won't be Microsoft — they just provided a tool, and that's legal. No, the defendant will be the downstream consumer who incorporated the code spat out by the LLM.
It doesn't matter whether the LLM "learned" — intent is irrelevant and the defendant will have committed copyright violation regardless. The LLM can't copyright-wash the code — nobody can.
The point is that there are repositories on Github that contain commits from non-Github-users. For example, not every single commit to the Linux kernel has been made by a Github user, but there is a mirror at https://github.com/torvalds/linux .
> You're playing a dumb game
For what it’s worth I’m not a lawyer and surely get some things wrong but licensing is an area where I have a certain amount of practical experience.
There’s some interesting ground to cover on this topic but we aren’t getting there and the incivility isn’t helping. Please consider this passage from the HN guidelines:
https://news.ycombinator.com/newsguidelines.html
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
Also, take it up with Torvalds:
"And yes, at least under US copyright law, and at least if you see Linux as a "collective work" (which is arguably the most straightforward reading og copyright law, but perhaps not the only one) I am actually the sole owner of copyright in the *collective* work of the Linux kernel."
If you believe Linus, and I do, you can see how your argument is devoid of logic when Linus himself owns the collective work and put it on GitHub.Don't forget the Office 365 "productivity" measures.
And the latest insanity, at least for me, ads in the settings of Windows 11. Never ever put ads in any configuration menu of the operating system. That clearly says, it's not my computer, its MS's and they try to squeeze every penny out they can get. So they aren't any different from Google, they even try to push you to the Cloud and then make mistakes a cloud enterprise shouldn't make
The Free PC people must be wondering what they did wrong.
Why not? I think they're great: they make more money for M$. If the users don't like it, too bad. Why should I care about them?
>That clearly says, it's not my computer, its MS's and they try to squeeze every penny out they can get.
That's right. It's not your computer, and you should just accept that and get used to it. You chose this path, by choosing MS as your OS provider.
>So they aren't any different from Google
No, they are different: Google doesn't control my PC's OS. The only PC OS they control is ChromeOS, which only runs on certain (usually low-end) hardware, and isn't normally used for any serious work.
> GitHub Copilot is trained on all languages that appear in public repositories
https://docs.github.com/en/copilot/overview-of-github-copilo...
I don’t mean to be dismissive but unless you did something like that or documented it in some way it’s hard to give too much weight to a “source: trust me” comment from a brand-new HN account.
Think about Google what you want but Google is never selling your data, it's far too valuable to sell. Advertisers never get to see any of it.
The decision which ad is seen by whom is handled entirely inside Google and your data never leaves them.
0: https://survey.stackoverflow.co/2023/#section-worked-with-vs...
To be fair, I haven’t tried it in like 15 years…
If that sounds like 15 years ago, it’s still true (high praise!) except it is more capable.
Personally I find it hard to believe that people use it as their default ides and write a lot of code in it.
but dunno
If something better than vsc comes along, I'll move to it
It’s really a consumer’s decision when it comes to Microsoft. Either you use the basic/free version in exchange for your privacy, or you pay for a business account.
I’m not sure Google really offers that as equitably across their offerings. They seem more focused on driving their ad/data collection business rather than private, enterprise solutions (just my two cents).
Do you think the amount of time you spend in the File menu is as monetizable as your search history?
I don’t even know how to find common ground with such a take. You’re saying vscode telemetry (which can be turned off by the way) is as bad as other data monetization because building a better product is profitable?
I hope I completely misunderstood you.
And why would Microsoft sell it? Like would you want to sell information like "X% of users can't figure out how to change the code formatting settings"? Or if you did learn something cool and useful, why would you sell it to the potential competition?
Why would anybody pay for customer traffic/browsing/shopping data for inside of a Walmart store? Nearly every retailer would pay for access to that, assuming the cost is reasonable. Vendors desperately want access to it.
Microsoft would be crazy to sell it openly. Its part of their competitive advantage, helping to protect $83 billion in operating income. They'd get tiny particles of revenue in comparison to the gigantic golden egg they're sitting on.
I can't imagine a lot of cases where telemetry would be on sale.
On one hand, telemetry is specific to a specific product. Knowing how often people use the File menu in Visual Studio doesn't help me with my own application that much.
On the other hand, telemetry is potentially a boon to the competition. Can you imagine if say, the LibreOffice project could buy MS Office telemetry to better decide what features to implement, or just to embarrass Microsoft by revealing how much something crashes?
For me, pretty much all telemetry seems to fall into one of those two boxes -- either useless because I'm not working on a clone, or useful and MS would be nuts to sell the info to me because I am working on a clone.
> Why would anybody pay for customer traffic/browsing/shopping data for inside of a Walmart store? Nearly every retailer would pay for access to that, assuming the cost is reasonable. Vendors desperately want access to it.
But Walmart would be crazy to sell it as-is, because why would they help their competition optimize their store layouts/warehousing? At best, I think Walmart might want to sell a derived service, like a slow trickle of recommendations to improve sales. Not the actual underlying data.
Microsoft knows the right CPM for these, because of their telemetry about start-menu usage. And while Microsoft isn't selling that information, per se, they are giving it to partners who consider running one of those ads.
Everything you do as a company, that creates value, also indirectly helps your ad prices if you start selling ads. That is the Google strategy: they used to deliver immense value with little monetization and now the balance has changed.
Having an issue with this with that reasoning is akin to having an issue with a company creating anything of value.
Any evidence to support this? Or did you mean everything you send which would be completely expected and normal per their terms? I see no IP traffic when typing but not sending data to ChatGPT.
I think you could only do this by caching the characters and mixing them in with later queries, but that behavior should be visible in the client side code.