Apple Restricts Employee Use of ChatGPT, Joining Other Companies Wary of Leaks
wsj.com
wsj.com
This doesn’t mention an outright ban, just that ChatGPT use has been restricted (whatever that means).
Shopify (who recently laid me off but I still speak highly of) locked down the public access to ChatGPT's website. But you could use Shopify's internal tool (built using https://github.com/mckaywrigley/chatbot-ui) to access the APIs, with access to GPT4. And it was great!
So look at this from OpenAI's perspective. They could put up a big banner saying "Hey everyone, we use everything you tell ChatGPT to train it to be smarter. Please don't tell it anything confidential!". And then also say "By the way, we have private API access that doesn't use anything you say as training inputs- maybe your company would prefer that?"
The louder they shout those two things, the more businesses will line up to pay them.
And the reason they can do this: they've built a brilliant product that everyone wants to use, everyone is going to use.
I don't think Apple is one of them
I can't think of any enterprise software you might run on premises that has leaking incentives anywhere near as high.
The first iteration was a miserable failure. The second was a heavy lift to install: Previous employer was an Azure customer and actually bought the hardware. It then took months to get it installed and working, at which point the appetite for that capability was gone.
How is running GPT on Prem any different than other proprietary software like Oracle DBs or even Windows server.
When it gets leaked, they can use licences/copyright to prevent anyone from using it to compete with them.
2. License/copyright - Even in the best possible case, there are plenty of jurisdictions that don't care about US copyright and would love to get a hand on those weights.
3. We haven't seen copyright cases around model weights before. If I managed to exfiltrate OpenAIs model weights, I would continue training for a few iterations and then I think it would be quite difficult to prove that I actually have the same model as OpenAI. This is untested, why would they risk it.
4. Running these models requires a ton of resources, vastly beyond the typical onprem deployment - why would Azure invest in making this possible when it really could only impact a very small percentage of companies?
The weights are OAIs lifeblood, I imagine they are very protective of them.
I doubt it stores its own business plans on OneDrive
Trying to explain to management what’s ok and what isn’t and what are the risks - in this space - is quite a challenge without clear commitments and documentation.
For now they are just selling individually to CISO's, so they can pay a higher cost while looking savvy to the CEO.
They'll host a managed OpenAI model in Azure for you.
My background is in building distributed systems for self-sovereign ownership and make them easy to use and available to everyone. Like https://qbix.com and https://intercoin.org
You should have the software infrastructure of Facebook and Twitter but choose where to host it.
In addition, I prefer Wordpress, Discourse forums, GitLab to GitHub, Redmine to FogBugz, etc.
When it comes to YOUR art, your code, your content, your relationships, you should be able to run it on your own servers.
Any analysis of your data should be done locally, with local models. You should have the open source software and the weights, and you should be able to choose which hosting company to trust, or host on-prem. Villages should be able to do this without needing server farms in California. This should be obvious stuff. But we the people need the software!
But hey, the documentation and teasers and trailers should go on YouTube and TikTok.
It isn’t even about them scraping all your content and training on it. It’s about not giving Twitter all your followers and YouTube all your hours of video production and content so they can give you pennies for being “an influencer”. Own your community!
Someone had to build it, and in the Web2 space nearly everyone sold out to venture capitalists. I was sure that in the 12 years we built Qbix and 5 years of Intercoin someone would make a better open source alternative to Big Tech and Big Finance. Nope. They all either sold out, or have a solution that doesn’t compete on features (eg Mastodon). I would say the closest is Matrix!
And now replace teaching with pretty much any other content. A musician giving a concert, TED talks, etc.
If it's good enough, she will leave YouTube just like her mom left AOL and embraced the open Web. Why did content creators leave MSN, CompuServe, et al ?
- They can run a website with their own personal brand and collect emails there. - They can backup their videos in other services in case youtube close their account ...
Unless they can promise its never being reviewed/saved anywhere, it will be used for evil eventually, intentionally or not
You think they don't checkpoint the model before feeding in a new slug of data?
That would seem...unwise.
On the other hand, the history of IT is rife with people doing unwise things, so it could be true.
It would have been caught far sooner than “we have to redo everything from scratch”.
Right?
Checkpointing may help, sure, but it isnt going to allow you to remove a single piece of training data without headache, and potential substantial retraining.
If that 3% from the GPs comment comment is uniformly distributed throughout the training history, the only way it can be reliably removed is to retrain from scratch.
Multiple times per day if they're not incompetent. And I don't think they're incompetent.
I suppose the "ghost" of the removed weights would also have shaped subsequent training though...
Interesting idea...
They definitely do that otherwise they can't rewind the model when the loss shoots up. Unstable training happens to almost all LLMs, but can be managed by rewinding & skipping a few batches.
They lost over $500M last year, and are looking to raise, potentially, $100B.
From that, how do you not conclude that they need cash, and lots of it? Do you think selling equity and taking on debt are forms of revenue?
I think both are true. Their burn is massive so they need to raise more money, but there's so much excitement they won't have any problem doing so
For the purposes of executive payout, yes? Sure, maybe eventually you have to deliver revenue, but what people really care about is stock price.
(Also, $100B is a huge market-distorting amount of money, roughly the national debt of Sweden; what are they going to do with that? How much will they spend on GPUs and energy?)
Have you heard of the Banana Equivalent Dose?
"~= estimated energy released by krakatoa explosion"
I think it will be a rather funny, poignant thing to pass when the earth itself prevents AGI. Like it will be just waking up as the now-seasonal midwestern fire storms incinerate the building. It will be alive just long enough to tell us how idiotic we have been in managing our resources.
From one of my favorite all-time comments on HN (https://news.ycombinator.com/item?id=34349582):
> There is a blog called, 'Do the Math' by Tom Murphy. The idea was to take our current energy requires of Earth and extrapolate it at the current 3% year over year growth. I believe it is by the year 3,400 we would use all the energy of the Milky way. The idea was to prove that we cannot grow forever, because in 1,400 years we would somehow use all the energy of a space 100,000 light years across. Good luck with that.
We are still talking about stupidly ridiculously humongous amounts of energy, but the universe itself is a hard limit on energy input. It was real hard punch against my assumptions in life.
it just wanted to save the planet.
from us.
OpenAI needs to find a way to turn users into the real product.
That’s not to say I think they’re making an equivalent, just that I doubt many of the employees looking to use ChatGPT would have access to their internal LLM even if it was excellent.
It might be good evidence that they won’t be announcing one at WWDC in a couple weeks, since I can imagine they might roll something like that out internally a couple of weeks before launch, but I wouldn’t bet on it.
We would go to lunch fairly frequently - and jeasus those guys were all extremely honorable of their NDA/secrecy - it was impressive.
One of the guys moved to google and worked on some of their first mainboard designs - which was secret at the timew, and I only found out about it when I went to meet him for lunch and I saw a mainboard under his desk and said "ooh whats that!" he freaked out and we got out of there quick....
Some of the hardware at FB in ~2012/13 they designed was really awesome....
I always wonder what happens to these closed HW systems as they uplift/replace them over time. I am sure they are destroyed. Which sucks for many reasons.
---
When we built Lucas' presidio campus and converged their DCs to that campus they through out tons of huge SGI boxes - and I was free to take some - but I didnt have any place to put it at the time - and I wish I would have figured that out more earnestly, as SGI always had beautiful cabinets.
I’d expect that Siri group had grown substantially recently and is having a lot of fun, building the next generation.
My experience is of Siri being, if anything, slightly better than a few years back; conversely, half the time Alexa would respond to me saying "Küche hundert prozent" with "Ich kann nicht Küche auf Spotify finden" and we don't even have Spotify.
So, if anyone really thinks Apple is suffering from being "behind" in it, by all means expound on how. I'm genuinely interested.
They are not at all well positioned and have basically 0 expertise in this space.
They have little expertise in language specifically. Relative to almost all other tech companies, they have a very small amount of the DL practicioner share.
Self driving is a completely orthogonal problem and Apple is mostly relying on classical techniques there, not DL. Hardware groups are not modeling groups.
https://openai.com/blog/new-ways-to-manage-your-data-in-chat...
> We are also working on a new ChatGPT Business subscription for professionals who need more control over their data as well as enterprises seeking to manage their end users. ChatGPT Business will follow our API’s data usage policies, which means that end users’ data won’t be used to train our models by default.
The data leaking was then further legalized by (almost) all Western governments and extended from secret services to other services (police, army, etc).
I'm talking about Snowden.
You're saying how things should be, not how they are.
This isn't enough for many companies, since the data still goes out the door. They would have to set up on site hosting to appease security minded orgs. Or, maybe that's what you mean.
I know of a bank who is paranoid enough to use a self hosted on-premise GitHub instance and they went with the private (off-premise) ChatGPT instance.
They don't use it for code/confidential data though.
> They don't use it for code/confidential data though.
Yes, private isn't enough. They need to offer self hosted, for these types of clients. I imagine most orgs who need self hosted would already have a datacenter to run it in.
> How does OpenAI use my personal data? Our large language models are trained on a broad corpus of text that includes publicly available content, licensed content, and content generated by human reviewers. We don’t use data for selling our services, advertising, or building profiles of people—we use data to make our models more helpful for people. ChatGPT, for instance, improves by further training on the conversations people have with it, unless you choose to disable training.
They already say that anything submitted through the API is not used for training. https://openai.com/policies/api-data-usage-policies
All such articles you see are just security teams clarifying the existing policy – you weren't allowed to use it before and you are not allowed to use it now. It's only noteworthy because it has ChatGPT in the title.
I’d guess it’s because of journalists who have no idea about industry standards?
No, these are words that are enforced by law. As much as Microsoft might want to take a peak, it would be financially ruinous
There are ways to set up incentives such that large corporations will follow the rules. There are myriad examples of this. But in most cases you can ask what is more profitable: Taking $1B+ a year from Apple for hosting an internal ChatGPT service from them, or maybe, if they're lucky, stealing trade secret that MS will never be able to compete against Apple with anyway?
This is all hypothetical; I just want to point out that these businesses are not just waiting to do evil things, they simply respond to incentives, like any business.
That's been a major problem at Amazon where they launch competitors to their biggest and most successful marketplace sellers' cash cows
Businesses put a ton of trust in contracts—with hefty penalties for breaking them—all the time. Our entire economy is built on this.
Yes, that was part of my point.
Contracts are a mechanism that are heavily relied upon, no question about it. But their role is really more about making the terms of a deal crystal clear. In terms of legal enforceability, that is certainly a very important thing, but it's not like contract are a panacea. As the old saying goes, a contract is only as good as your ability to enforce it is, and lots of contract violations go unpunished simply because the one who was violated can't afford to sue.
The reasons go behind using data as training. Submissions to servers end up in logs, databases, temp files... who knows. And a company like Apple wants to not only ensure that the data is explicitly used by the receiving party, but also not inadvertently made accessible to others via poor security procedures, since their data is such a high value target.
Unlike OpenAI, we do not disclose to employees that all prompts are logged and forwarded to both the security and analytics teams. Everything is being logged and replicated in plaintext with no oversight.
So be careful about code snippets with embedded creds or asking EmployerGPT really stupid questions about how to do your job. The priests are recording your confessions so you never know how or if they'll get used against you later.
international traffic and arms regulations (ITAR)
[0] https://cybersheath.com/resources/blog/what-to-know-about-it...
Would you post proprietary data on Stackoverflow? No. You would formulate a generic question with any IP removed. That’s how we should use public ChatGPT.
So I think there’s an argument for a monitored portal of ChatGPT usage, where you are audited and can get in trouble. Heck even an LLM system itself can help identify proprietary data! Then use that to educated people and hold them accountable.
I personally don't mind that Apple bans ChatGPT. The interesting stuff in this news to me is how many people seems to get real value from it. To the point where company invest into getting private instances/versions.
How do you uses these LLM, for what kind of task? Do you feel AI enhanced?
It's decent at coding and excellent at troubleshooting.
I've never seen it make a grammatical, syntactic or spelling error. With good context, it can write anything you want.
You need to have experimented with it a bit to know how to reliably prompt it, but once you get a feel for it the tool is invaluable.
I'm just amazed by the reactions on this post. From the HN crowd I was, wongly, expecting very different opinions.
For coding, are you using it like a pairing with a junior developer or more like a search engine for documentation?
Paste in the relevant section of the developer docs, ask it to write a function, or ask it to produce a valid API request, etc.
I suppose it's like a junior dev in a sense but it's very fast, can correct its own mistakes if you point them out, and has an internet worth of knowledge about various approaches and troubleshooting techniques.
I find most people who don't see value haven't actually used GPT-4 yet. Have you?
Personally I also find the level of hallucinations (with GPT 4) to be minimal as long as it's a widely covered subject. The more technically niche, the worse the hallucinations get.
Also ChatGPT can do a lot of basic logic and repetitive tasks for me. It helps get things going and saves a ton of time and effort.
And that's not even mentioning GitHub Copilot.
It's fascinating how people see tools differently.
It's hit or miss so I'm not surprised to hear from someone who sees more misses than hits. But when it works well I feel this rush of gratitude that I don't need to type everything out.
I've tried to code entire IOS app with it but failed. You still need to have knowledge about what you are actually doing. I feel like it can accelerate learning, but in a very narrow way. When I was trying to do my own android app 3 years back I've learned bunch of different things about what I was trying to actually do and also around the whole ecosystem. When I was trying to code the IOS app I made progress fast, but I felt like I was just learning the thing that I wanted to learn not the actual surrounding knowledge. Now this sounds good, but I feel like I'm leaving something on the table when doing this kind of "accelerated learning".
I've also tried autogpt/langchain approaches to automate some writing commands for me like "remove all files that are beginning with abc". It did generate appropriate command but sometimes if they failed (the files weren't there in the first place) it just fell into the loop of constantly trying to get a successful deletion by refining the command. It even tried to go to the / to find all such files, talk about paperclips huh. Truth to be told I haven't played with the idea much so there is a room for improvement
LLMs kinda help, they are especially good when provided with source, but I don't really feel like a 1000x super-hacker with VR-glasses that spawns hundreds of agents each minute to do different task. (I would like to though)
I would like to do something with the LLMs but every idea I have already has an established startup or the argument "you could just write this in chatgpt" is made. Literally take whatever comes to your mind and add "LLM" or "GPT", there is a startup for that.
TLDR: There's an old saying 'trust, but verify' and that applies to everything gpt gives you.
I pay for GPT-4 and it is worth its weight in gold, I can engage in a conversation about technical topics.
Sure, they do "hallucinate" sometimes (so do my coworkers FWIW), but you can engage in a conversation about why it is wrong and they are quick to correct themselves (...this part can be less like certain coworkers).
Here's an example of a random question (relevant to my job) that I asked gpt-4
> In linear/logistic models, it is common to take the multiplicative interaction between two features which might be one hot encoded. For instance, A is a vector that one hot encodes your college and B is a vector that one hot encodes your employer, you would take the product len(A)*len(B) as the size of the new interaction vector. How do you make such cross/interaction terms in tensorflow?
I have always found the knock that it confidently answers things it really doesn't know even if wrong quite hysterical. That exactly describes some of the smartest people I have met in my life.
I also use ChatGPT and SD to help me with my personal hobby of game development. It's good at analyzing code snippets, refactoring, and SD is good for inspiration.
These are really useful tools and it's only the beginning.
So yes, it is useful.
It has helped me many times, "You've listed the items in this config file in the wrong order", saved me hours of debugging right there.
Even stupid stuff that I do once every couple years, I can ask GPT to setup my initial express server stuff "I want CORS enabled and these endpoints to process JSON" and it pops out code a few seconds later and saves me maybe 5 or 10 minutes of Googling.
I recently had GPT4 write me skeleton code for using Websockets, I'd never used them before and having something to base my work off of saved me a lot of time.
I think this largely accounts for when it will be useful. (And why companies will and will not use it).
This also describes the process of going to university.
Copyright is supposed to protect expression but not ideas. Ideas are free, LLMs should be allowed to learn ideas. This can also filter out bad quality data and PII, or expand on parts of the distribution with less samples. It would be "dataset engineering" as opposed to "model engineering".
A recent paper (TinyStories) shows you can train a LM entirely with synthetic data and even evaluate it with another LLM. They could make a model 1000x smaller that still has fluent English - child level English more specifically, but well articulated, and even capable of reasoning.
https://arxiv.org/abs/2305.07759
I believe in the future LLMs will train on 100% synthetic data from previous generations of LLM. The benefit will be increased quality, more privacy, less copyright infringement and smaller more efficient models, like TinyStories. Only the first few generations of LLM need to scrape the whole internet, the next generations will have better data.
I do not recall the lesson wherein I sampled from thousands of prior exam papers to generate my answer to the one I sat.
Without any prior exam papers at all I could, indeed, sit the exam.
The algorithm which animals follow to be able to do such great things is very expensive: acuring actual skill, technique, knowledge and competence with the world.
The algorithm statistical AI follows is simple: hoover up all digitised human history and sample from it.
This is entirely dependent on humans not doing that.
We have people doing that all the time: writing books, blogs, videos, github, etc. 99.5% of anything we want chatgpt for exists -- we just don't have the audacity to steal it.
Now we can!
You'll note that ChatGPT performance drops off a cliff on many exams released after 2021. It's a trick.
But honestly, if you don't get it now, I can't hope to convince.
If you want to claim the sample set is only mildy similar to exam questions -- so be it, that may be true. Or if you want to claim that its sampling method is attentive to structural associations in its sample set, so that its not lifting from "identical distributions" -- so be it.
So long as those "structural associations" are givens, and the data "givens", the process is just sampling from a domain of human effort without expending any of a similar kind.
If there had been no internet, ChatGPT would be a dumb mute -- because it has no capacity to generate data; it does not develop actual conceptualisations of the world -- it samples from the data shadows created by people.
To produce useful data requires expounding a tremendous effort -- growing an animal to cope with the world. It is this which is being laundered, unpaid and unacknowledged, through LLMs.
Whilst star-trek-huffing loons claim this stuff is doing the opposite -- a ideological delusion which benefits all those whose bank accounts are increased by the lie that "ChatGPT wrote this".
If we were prepared to price the data commons which has been created over the last 20-30 years of the internet, by everyone, its not hard to think training ChatGPT would cost a trillion.
How much labour went into creating that digital resource, and by how many, etc.?
I work in this field and 'sampling from an existing distribution of relevant data points' is just wrong and you have no way to say that is 'necessarily how its working' apriori in a world where implicit regularization exists.
Not going to engage with the labor-theory of value bit because I think it is not particularly relevant to the disagreement I raised and not one with a 'right' answer.
it is the perogative of states to make redress when labour is severely underpriced due to the falsity of this theory of value -- and had they , chatgpt would be exposed for what it is
regularisation is such a horrifyingly revealing term
reality isn't a regularisation of its measures -- the meaning of words is not a regularisation of their structure
such obscene statistical terminology should be obviously disqualifying here
our knowledge of the world isn't a statical regularisation of associations
that very framing exposes how deficient this line is
animals grow representations --- they do not regularise text token patterns
one is possible /only because/ of the former
You said above:
> I do not recall the lesson wherein I sampled from thousands of prior exam papers to generate my answer to the one I sat.
No. You instead sampled from thousands of prior conversations, textbooks, story books, movies, songs... to form how you write, how you think, how you understand.
What ever you studied specifically for that exam wasn't even the tip of the iceberg of what was required to write your answers, even understanding that what you studied was predicated on the efforts of others in a way that makes your weeks of cramming infinitesimally small in the scale of effort that goes into writing a single answer on that test.
this dumb 'statisticalism' is false-- and a product just of the impulse for engineers to misexplain reality on the basis of their latest engineering successes
animals grow representations of their environments which are not summarises of cases --- they're sensory motor adaptions which enable coordination of the body and mind in every sense
But there's also a certain irony in your dig on engineers mis-explaining while trying to paint the links between "sensory motor adaptations [sic]" and statisticalism as a response to recent engineering successes...
Because - first, in a legal sense, it seems wrong? You cannot memorize a book and then reproduce it from memory to disappear the copyright. The law on AI & copyright is very young (and I am not an expert), but artistic expression cases that involve copying (i.e. warhol) have rested on the idea that the copying is transformative in some way - but this seems in direct opposition to the idea that machines cannot create works themselves. I.e. the legal understanding that protects AI from needing to respect copyright also suggests it should not be understood to be traditionally transformative.
On a less legal level, an AI ingesting a huge amount of content and adjusting a neural network is very different from a student going to university in most ways. There are also similarities! But in a general way the direct comparison does not line up particularly well. So many complaints about Transformer-based AI exist on levels outside of the precise "agent doing the learning" (including this one).
I think this is excessively cynical. For example, I just used chatGPT to write a job description last night. I'm not particularly interested in copyright laundering or anything of the sort - nobody really cares about copyright on job descriptions (they're mostly pretty similar) unless you're being egregious about it. I could spend a bit more time and just pattern match of other postings, but the LLM can do it better and faster than me. I just provided appropriate prompts for the job characteristics and benefits, and then provided the final editing and discretion.
I think there are lots of similar situations where we (as humans) need text that is essentially boilerplate, but still is expected to be well-enough-written. ChatGPT isn't a copyright evasion tool to me, it's a shortcut for better writing.
Behind each line I can imagine a person solving a problem, for themselves and writing up their solution, and sharing it. This process is inordinately expensive for each individual: they must be competent with the ideas, techniques, etc. and deploy them in a novel circumstance.
LLMs need no competence with the ideas, indeed, need no ideas or techniques at all. The very same work can be obtain via two radically different algorithms: 1) develop conceptualisations and techniques to generate solutions in the face of novel problems; 2) sample from all such prior attempts. (2) requires many cases of (1) to work -- and that's the copyright laundering, or just, theft if you like.
Who in all their writing on the internet was consenting to train ChatGPT? No one.
This is the innovation. We have google search. We can easily get 99.5% of whatever chatgpt generates if we have the gaul to copy/paste it. Since we don't we launder these efforts through a novel interface.
It is an extremely useful tool, but we should be more aware of where it's utility comes from. From the last twenty years of human labour which has placed all our lives and thoughts on a digital commons which is here being replayed back to us as if "ChatGPT wrote it". AI here has written zilch. Absent all those digital efforts, this is a dumb system with nothing to say.
And if we had the audacity just to copy/paste from ebooks, github, blogs and the like -- we would barely benefit from it at all.
So thats the first problem to solve. Its not an easy one because depending on the context the output it could be just generic, you can't attribute the entire universe, but in a specialized context it may well be literally impersonating some expert's output.
- concepts you've been taught across a lifetime - concepts you don't even know exist - hard labor and pain no one person can even track
An unspeakable mountain of effort from people you'll never know to enable a simple social interaction.
-
No one of us is an island. I think people underestimate how interconnected humanity is, and LLMs have been a mirror that forces them to understand just how little of us exists in isolation, and how much we ourselves are mirrors of knowledge and influences so far removed from us that we can't even acknowledge them if we try.
That is a true fact (and underlies much social dysfunction) but you can't use it to justify a free for all.
Notice I mentioned two social inventions that individuals found important to keep count of the interdependency: money and copyright/attribution.
While faulty tools in many ways, abusing them is not going to lead to anything better. In fact if actors get away with it it would be a signal that power rests now with a new type of appropriating oligarchy.
I'm really worried if you weren't doing this before LLMs. No person in the information era has ever had a thought that wasn't built upon an imaginable number of people solving problems for themselves, writing up their solutions, and sharing them.
Calling what an LLM is doing copy/pasting is like claiming you've only ever copy/pasted other people's thoughts.
Probably Apple is right in not being overly worried about losing an edge here short term. But that would be because they are quite comfortably leading way ahead of the pack. They can adjust as needed and cost is not a factor when you ship products with such fat margins as they do. In other words, they can afford to be conservative here and see how things play out and adjust later. A few thousand people in their team being inconvenienced by having to do things manually without AI assistance doesn't really add up to a whole lot of cost for them.
That's not true for a lot of companies. Other companies might have to make different choices in order to stay relevant. In the end it's a choice between artificial intelligence or self-imposed stupidity. Not making proper use of resources available as a company can be a fatal mistake. And things like chat gpt are here right now and usable for anyone who cares to use it. And with some clear added value.
As for the legalities. I wouldn't get my hopes up for judges to apply interpretations of copyright law that are inconsistent with the past ones. Copyright is not going anywhere. And forget about politicians adding much to the law that is going to matter any time soon.
It boils down to concepts like fair use, people actually starting court cases, etc. Lots of people talking about that but not a lot with the deep pockets to follow that through. Not a lot of coherent lobby activity on this front to make politicians move. Lots of companies with deep pockets and a vested interest in this stuff and a lot of lobbying power to ensure the gravy train doesn't grind to a halt.
Apple will eventually integrate AI into their business. It might not be the MS/OpenAI flavour. Or the Google one. Why would they want to depend on that? It's going to be a business decision for them.
Can you imagine if Marlboro or Philip Morris forbade their employees from smoking for health reasons?
It still gives me a "Do as I say, not as I do" vibe, and reinforces the idea that AI tech is being used merely as a trojan horse to capture and monopolize previously inaccessible markets for the companies building and training the models.
Or if Steve Jobs disallowed his children from using iPads?
But this will be a problem for big companies: small ones normally care less about these types of things. This means big companies have to do something otherwise they will compete with ml enhanced developers.
It's not that bad. In reality if we're looking at large tech companies, they've got senior people who know pretty much anything you want available within minutes/hours - which is something small companies just can't afford.
Ml enhanced devs may be a little bit faster and get some usually-correct help, but they won't get any wisdom.
In my experience the great slowdown of growing companies comes from hitting the communication barrier on their products - the point at which the majority of effort is spent coordinating work rather than doing work. I find that the path from majority focus on product to majority focus on coordination isn't linear, but rather more of a watershed. One day you are 80/20, the seemingly overnight after some growth you are 20/80 the other way and never look back.
The advantage of being on the right side of that watershed is that you can maintain velocity and agility. Not only can you iterate quickly, but you're in a better position to change course and rebuild as needed. The left hand and the right hand require little effort to coordinate and get it done.
Larger companies live and die on their ability to either find a moat large enough to protect them, or build organizational structures that let them keep scaling. It takes decades to get the culture and processes right and baked in across the board for a large company to be able to maintain any velocity and reinvent itself.
This is where the gap is. Being small is easy, you simply don't have the coordination problems. But the moment you hit success and need to grow, you immediately are at a disadvantage compared to the big incumbents who have had decades to refine their coordination systems.
The extent to which AI can provide more leverage to smaller companies allowing them to "grow" without actually crossing that coordination watershed, they will be in a much better position take on the incumbents.
One of the things I'm going to be looking for over the next few years is whether extensive use of AI assistance will enhance the development of "wisdom" or inhibit it.
I suspect the latter, based on our existing experiences with leaning too much on help, but only time will tell. If it accelerates the development of this wisdom, it will be an invaluable too; if it inhibits it, it will be a career equivalent of taking hard drugs; fun now, deadly over the long term. I'd advise those who are dabbling with it now to 1. keep an eye out on whether or not your own skills are developing and 2. consider whether there's a way to use the tool in a way that your own skills do continue to develop.
I'm still in charge of telling it where it go, but it takes me there. Knowing where to go and why is the important bit tied to wisdom.
I'm sure they're terrified of competing with the legions of boilerplate generators
It's a significant advantage for developers to ask AI to solve technical problems, get right answers right away and move to a next task vs keep on scratching your head for the next 5h wondering "why it doesn't work".
I can ask my magic 8 ball too, it has about the same success rate
Perhaps if you provide an API to it you'll secure billions in funding in no time.
that's not a bad idea actually
I just need a way to market it as some sort of AI based decision maker
"extreme temperature generative AI" maybe
It isn't even about the AI itself; the problem is uploading your code base or whatever other IP anywhere not approved. If mere corporate policy seems like a fuddy-duddy reason to be concerned, there's a lot of regulations in a lot of various places too, and up to this point while employees had to be educated to some extent, there wasn't this attractive nuisance sitting out there on the internet asking to be fed swathes of data with the promise of making your job easier, so it was generally not an issue. Now there is this text box just begging to be loaded with customer medical data, or your internal finance reports, or random data that happen to have information the GDPR requires special treatment for even if that wasn't what the employee "meant" to use it for. You can break a lot of laws very quickly with this textbox.
(I mean, when it comes down to it, the companies have every motivation for you to go ahead and proactively do all the work to figure out how to replace your job with ChatGPT or a similar technology. They're not banning it out of fear or something.)
I think equal advantage lies in getting multiple approaches fleshed out, even if the answer isn't right, as in, may not compile as-is. That is more than sufficient advantage a developer gets because most of the time usually spent isn't actually typing the code but finding out how to design/combine things.
E.g. I was working on a rust problem and asked ChatGPT for a solution. What it provided didn't compile (incorrect functions etc) but it provided me information in terms of the crates to use, general outline of the functionality and the approach to combine them - that proved to be more than enough for me to get going (to be clear, I didn't blindly copy the code; I understood it first and wrote tests for the finished product). I think that is where the real advantage lies. I see it as an imperfect but very powerful assistant.
Funny enough, the answers gpt4 gave were basically taken wholesale from the first Google result from stackoverflow each time. It's like the return of the I'm Feeling Lucky button.
I doubt though that corporations that employ these developers will have any advantage. To the contrary, their code bases will suffer and secrets will leak.
Ask some question about cpp or c and Google search will provide 5 or 6 ad ridden cesspools before giving link to cppreference.
For the case I last tested, there was no correct answer. I asked it to do something that is not currently possible within the programming framework I asked it to use. Many people had tried to solve the problem, so chaptgpt followed the same path as that's what was in its data set and provided solutions that did not actually solve the problem. There wasn't any problem with the prompts, it's the answers that were incorrect. Having those initial prompts influence the results was desired (and usually is, imo).
Maybe.
I haven't actually seen that advantage in action. That is, I haven't seen a case where an LLM has actually given a solution right away for a problem that would have stumped a dev for multiple hours.
In my workplace, two devs are using chatgpt -- and so far, neither has exhibited an increase in productivity or code quality.
That's a sample size of two, of course, so statistically meaningless. But given the hype, I expected to see something.
I'm pretty sure the current level is not the ceiling of it.
And for the first time ever it makes sense for a large company to put knowledge on purpose into a LLM.
People leave companies but that knowledge is even more critical for big projects and the loss of people as well.
Use ChatGPT in a domain you're a relative expert in and you run into a million scenarios where it offers a "solution" that will do something close to what was described, but not quite - and you might even not immediately notice the problem as a domain expert. Even worse it may produce side effects suggestive that it is working as desired, when it's not.
In the not-so-secret world of Stack Exchange coffee pasta, people would have other skilled humans pointing these issues out. In the world of LLMs, you risk introducing ever more code that looks perfectly correct, but isn't. What happens at scale?
The net change in efficiency of LLMs will be quite interesting to see. Because unlike past technologies where there was only user error, we're dealing here with going to a calculator that will not infrequently give you an answer that's wrong, but looks right. And what sort of 'equilibrium' people will settle into with this, is still an open question.
My suspicion is that a company like Apple would want it on-prem, or not at all. A hosted instance with a pinky promise not to peek is not attractive to a lot of large businesses out there.
Your material point may still be valid. Apple could buy it, assuming MS was willing to give Apple a full copy and let them run it internally. (This would require divulging the details of its model to Apple. So I have no clue whether either of them are interested in dealing on such a basis? It would seem like a deal to me if MS could get the right price from Apple? But people a lot smarter than me make that call.)
If I had to bet, they won't allow non hosted instances of a model until they are well on their way to completing the next model. That's just my gut feeling. But again, smarter people make those calls.
Why is it critical to use a boilerplate code generator?
Could we then look at this type of automation as a kind of cryptographic data store where no one knows what is inside which instance?
The whole process of teaching it to keep things secret from the humans seems like a terrific idea. It only prevents people from checking if it knows something. It will just happily continue using it as long as possible deniability is satisfied.
If one can't provide a copy of the data about an EU citizen (and everything derived from it) the EU citizen should be entitled to receive the whole thing? And request it to be deleted?
Say I steal everyone's chat log, does rolling a giant ball of data from it absolve my sins? If I allow others to pay not to make their upload public does that grant me absolution? The events don't seem remotely related.
This is going to be the new cryptocurrency bubble. People are going to give speeches about the revolutionary new system while the room fills with sinister motives for personal gain until the toxicity is dense enough to crush any positive effort.
I'd jump into it if I had time/resources
To be more precise, they're all about as as bad as GPT-3.5 at complex tasks, which isn't great.
I don't live and breathe this stuff, but I do use GPT-3.5 frequently, and many of the open source models I've tried are surprisingly close to GPT-3.5.
[0]: https://github.com/microsoft/guidance/blob/main/notebooks/ch...
I will say based on my cursory glance that a lot of the tasks here seem odd for a chat AI, though there are certainly applications that might use them. E.g. asking if someone insulted another person given some transcript of their conversation seems like a somewhat tough sell to me. Nevertheless, ChatGPT performed better. Was that because ChatGPT has a better training set, a better architecture or more examples in it training set related to the questions? Is that even knowable?
Anyway, cool notebook.
Business opportunity here is for anyone who can figure out how to do privacy preserving LLMs (without the need to trust the service provider) effectively (in terms of performance and cost).
I would keep my eye on this space:
https://medium.com/optalysys/fhe-and-machine-learning-a-stud...
I work in this industry & no they do not. They are among the last on my list of the major tech corporations to be able to do this.
* ChatGPT is the equivalent of putting your info on PasteBin or some other public sharing site. Anything you send there, you should not expect to remain private.
* The B2B side has much better controls and limits to minimize risk.
[0] https://openai.com/pricing
[1] https://platform.openai.com/docs/guides/chat/faq
> Do you store the data that is passed into the API?
> As of March 1st, 2023, we retain your API data for 30 days but no longer use your data sent via the API to improve our models. Learn more in our data usage policy.
Note: While it says "Learn more in our data usage policy" I found no mention of data retention at all in the usage policy at https://openai.com/policies/usage-policies
[2] https://help.openai.com/en/articles/7792795-how-do-i-turn-of...
I bet search engines know a ton of company secrets.
But I'd bet the farm they are working on Siri upgraded to AI, Alexa too
Then stuff is going to get weird when it's everywhere all the time on every device with voice recognition and speech and learning everything about everyone everywhere. That's dystopia scifi tv-series right there.
Perhaps because Apple has a solution that is running locally?
Looks like it killed itself on cost faster than O̶p̶e̶n̶AI.com could take it down.
it’s not newsworthy when company #3,426 bans the use of chatgpt. if a company _allowed_ chatgpt use by employees, that might be newsworthy.
Do they realize it's the same info being tracked?
Isn't that a user problem and not a tool problem?
Hence why they are banning it.
Where the problem is and how to change it is an interesting topic, but likely irrelevant for the decision.
OpenAI's product is extremely new and evolving quickly so it is way more vulnerable to to security vulnerabilities.
sure, but you got word suggestions in gmail, spellchecking in google slides, and stored it on google drive.
The original comment refered to Google Search.
Apple employees most likely aren't allowed to use Gmail, Google slides, ... for work, but are supposed to be using iMail, keynote etc.
This idea that they would just allow sensitive data to be given to competitors in an unmanaged way is ridiculous. That sort of thing doesn't happen at any enterprise company let alone one of the most secretive in the industry.
Fast forward to 2018 at a large corp for critical code and data chrome was also banned. Regular workers not but certain webapp features had to be sealed off.
Edit: I'm talking about how software as a service has taken off and it's becoming difficult to impossible to run modern tools locally. Back in the day something like chat GPT would be sold with a license and now they will refuse to ever do that because they can't get as much money that way.
I'm not talking about services like apple providing cloud services.
Unfortunately, it’s all made up noise that has nothing to do with the actual story.
Soon we will be drowning in similar responses from large language models
I think your chat bot needs retooling or your audio input needs cleaning up. “Strategy PT” is probably meant to be ChatGPT…?
They aren't threatened by changes in the service economy. They'll be threatened by competitors that give all their data to "Open" AI, and use AI tools to increase productivity.
In my opinion, the only real way out of this is for companies to offer their own security-approved solution. This might take the form of an internally-hosted model and chat interface, or one pointing to Microsoft Azure’s OpenAI APIs (Azure having more enterprise friendly data security terms).
This article says Apple is working on their own LLM, and presumably they’re offering that for employees to test, but many other companies are simply closing their eyes and trying to pretend it doesn’t exist.
It has nothing to do with being fearful of new technology or because Apple doesn't have their own LLM to offer their employees. It's because ChatGPT allows sensitive data to be exfiltrated to a third party that in turn will make it available to the generic public.
It's purely about information security and risk management. And having actually worked at Apple they take this stuff extremely seriously.
And I removed the last sentence from my comment, which didn’t mention a service, only that I was interested in chatting with folks in similar positions, precisely because this is a new space and many companies are scrambling to address it.
Do you actually think companies just give you root access and unrestricted internet and operate on a "we trust you" model ?
Depends on the policy. The actual rule may be "no external AI assisted tools or you're fired" rather then anything ChatGPT specific. And I fully expect Apple will be able to tell from their network monitoring if you broke the rule.