Google warns its own employees: Do not use code generated by Bard
theregister.com
theregister.com
The copyright of the output of these tools is not yet determined, it's a risk.
The advice not to use the tools in a risky way not only directly mitigates the risk but like an "employees must wash hands" sign in a restaurant restroom, it can be seen as a reasonable step that transfers some of the liability to the staff member.
It's not a verdict on the output of the tools, it's a reflection of risk aversion in a monorepo company that is very protective of what is in the monorepo.
But exactly, this is a hygiene thing. Staff will still be using these tools anyway.
The tainted code can become a critical dependency. It can get copy-pasted elsewhere. An engineer could look at the code before writing their own (obviously not illegal per se but makes it harder to repeal bullshit claims later). An engineer can just write similar code and then have no way to prove they didn't look (ditto).
I have no experience working with a monorepo but I've read Google has "an army" of people maintaining the monorepo. I would have imagined with enough support like this, you wouldn't need to copy paste code everywhere in a monorepo?
I'd appreciate any insight anyone can share on monorepo. I am fascinated by the idea but I've never had a chance to work with it and people who I've met who have either don't want to or can't talk about their experience with it.
Remember: swift fingertips sink ships.
How would he know it doesn’t understand if he hasn’t used it?
In fact, we were literally asked to.
That wasn't said, but the goalpost keeps moving in this thread.
The other one I'm thinking of; I know this isn't the case with Google because their repo is too big, but, if a piece of copyrighted code would end up in an average repository, it would quickly be distributed to all users that have a copy of that repo, so they can't take it back again.
You're talking about a company which happily pays billions in fines because its just a fraction of what they otherwise manage to get away with. They're not at all afraid of doing a lot of illegal stuff, I find it highly unlikely they'd be afraid to step into a minor grey area.
But that doesn't mean that the company will tolerate widespread illegality or allow systemic risks to build up. It's simply not how organisations do things especially those are highly public. Poor reputation doesn't just manifest in lower sales but also introduces unnecessary regulatory risks.
Being able to say to their (probably then ex-) employees "I told you so" will gain them nothing. There will not be nearly enough to get from the breaching ex-employee to compensate for any damages.
If the restaurant has signs like that up, and consistent messaging to their employees, and new employee training, and other "best practices", then I may still be able to sue them and win. But if they don't have all that stuff, I may be able to sue them for treble damages because of negligence.
That is, such measures may not remove liability, but it limits it.
Note well: IANAL. Others who know more, feel free to offer corrections.
I think we're already past the point of no return. Almost every active codebase in the world (at least the JS part of it) has been tainted by LLM generated code at this point if dependencies count.
If courts decided the output of LLMs trained on GPL code was subject to GPL, all the active code in the world would need to be released - which seems impossible to enforce.
As an employee, they take zero of the business risk, as no employee is liable for any of business's risks. It's also why CEOs don't go to jail if they're performing their duties, even if the company is liable for damages or commits a crime.
And this is the way it should be.
Where are you getting any of that from the linked article?
The definitive definition of irony o_O
This reminds me of two pictures that were circulating on social media a while ago. The first was Zuckerberg holding a book while billions of people are wasting their time in endless scrolling in the Metas; the second was that Turkish chef known as Burak eating a healthy lite salad.
The bottom line is that we sell stuff that we are not convinced to consume.
In fact i would recommend every dev to use their own stuff, "dogfooding" helps a lot to improve the product.
Imagine Google Bard outputs a piece of code for a major new feature, making them a shit ton of money. It’s then found out that code was actually taken copied, almost directly from average Joe’s hobby app, and now you’ve got a situation on your hands. Amazing for the average joe who’s expected to get a big pay day if the courts ruled in his favour but less so for the business owner.
Frankly, I can’t wait to see what happens when all copyright is out of the window. Obviously within reason, but wait until you see the innovation that comes about when the gates are open. The free-for-all of ideas is a fun time to live in that’s for sure.
Oh, so you don't really mean "when all copyright is out of the window", I guess. So.... how are we going to define "within reason"?
What that looks like, who knows.
Is anyone's productivity really higher if we have to spend all our time correctness checking the output?
Find someone at the same level of competency as you, set a goal, that person gets access to any LLM they want, you have to do it with no assistance other than reading stack or googling. And if you think your ego can handle it, find someone who you know is below you in competency (but not junior) and set the same goal and see if they surpass your own productivity or progress.
We’d have to set up some controls but you see where I’m getting at. It would be an interesting experiment either way.
It’s all in how you use the tools. If you’re zero shotting and hoping for the best then yeah, not great. But that’s where your domain expertise is supposed to come in.
WHERE Column1 <> Column2
The suggested improvement: WHERE Column1 <> Column2
When I pointed out that it just gave me exactly the same code, it apologised and suggested this improvement instead: WHERE Column1 > Column2
I'm trying to understand how this helps me become more productive...LLMs writing code have strange performance characteristics. Sometimes they produce working output many times faster than the fastest of senior developers, other times they can't understand a bug even when you point it out to them.
Try giving your LLM this prompt:
Please write a python script that loads a file named input.png, converts it to greyscale, plots a chart of the intensity of the pixels of the first row, and saves the chart as output.png . The chart should be a line chart with a blue line, and the lowest and highest points on the chart should be marked with a red cross. Thanks!
So LLMs can sometimes produce working code very fast. Other times, as you've identified, they produce completely broken results.
It gets the rough outline of a solution OK but very easily gets lost if the problem requires 2 or 3 levels of what I guess I'd call nested analysis. Or rather, it's fine for creating a surface solution, but it gets lost in the weeds quite easily for things like type system or borrow checker problems.
Because it doesn't reason -- it (re)produces patterns.
well that's one way to make code faster, just require the underlying hardware to compute things for you
But every time it gave me a very sensible solution.. Usually not something that worked right out of the box but never something totally stupid.
The last time I read the Github CoPilot agreement for example it used specific language such as: "Suggestions"; and seemed to indicate that copilot was merely inspiring developers through suggestion rather than creating code that could (or should) be copied verbatim.
I think in reality most people are copying code verbatim and not using AI as inspiration, but it seems like (legally speaking) that it is the safest approach to treat AI generated code as non-copyable..
That notion is consistent with what Google is stating here.
There are ways for developers to use chatbots that doesn't involve copy-pasting code into your own codebase. Just like there is a difference between not using stackoverflow (which is foolish) and not copy-pasting from stackoverflow (which is sensible).
I have used ChatGPT a few times as a developer and found it useful, but never directly copied code from it. I had it document my code for me though, so maybe if you count comments as code, I did.
At its core, an advertisement is made of two components: the subject - a store, a brand, a product, etc. - with which the victim is supposed to develop positive associations, ideally strong enough to motivate a purchase and/or advertising the subject to their acquaintances, and the message, which is meant to create those associations. Only the first part, the subject, has to be factual. The message does not, and in fact it usually isn't - manipulative bullshit performs much, much better, and generally the optimum for advertising seems to be asymptotically close to the line past which it would be legally fraud.
The only factual part, the subject, needs accurate handling, and is best suited for classical database systems - which is exactly how Google, and everyone else, is handling it. The message part - that's a good match to LLMs, which excel at producing convincingly sounding bullshit. For advertising, it seems what you need is to crank up the hallucinations a bit, but have some plausible deniability built into the whole system, so that when the LLM hallucinates in too obvious a way, no one can actually be held responsible.
It’s flat out lying, yet folks are using its output in arguments.
It's always good to be aware of how things affect human hormones and absolutely worth study and regulation, but it's a huge jump from frogs are affected to humans are affected. I know right wing people are world class long jumpers so this kind of "it could be!" really isn't helping towards a reasonable debate.
Barring any AI court case setting a precedent, We're still at 'AI generated works are public domain' aren't we?
There's probably going to be some big money arguing for the right to own content generated by their AI. I can't really see that being in the public interest though.
I didn't hear we were ever there?
https://builtin.com/artificial-intelligence/ai-copyright
"Can AI Art Be Copyrighted? It has long been the posture of the U.S. Copyright Office that there is no copyright protection for works created by non-humans, including machines. Therefore, the product of a generative AI model cannot be copyrighted."
Is that last bit from the Copyright Office, or is it the author's interpretation? Because I could just as easily imagine the battle being over who or what was the actual creator of the content (i.e. is it a derivative work), rather than whether the thing the machine created is eligible for copyright.
To remove the AI from the equation for a second, imagine that I took four images of living artists' work, placed them in a 2x2 grid, and called that a new artwork. There are two seperate questions to consider: (1) have I infringed upon the original authors' copyrights, and (2) is the new thing I have created eligible for copyright.
The stance that there is "no copyright protection for works created by non-humans" only addresses the second question, not the first question of how it interacts with existing copyrights.
Just like you don't grant copyright on a PDF to Adobe Photoshop... you grant it to the human who guided the program. Photoshop is a tool and the human is the creator.
https://www.theregister.com/2023/03/16/ai_art_copyright_usco...
> "For example, when an AI technology receives solely a prompt from a human and produces complex written, visual, or musical works in response, the 'traditional elements of authorship' are determined and executed by the technology – not the human user.
> "Instead, these prompts function more like instructions to a commissioned artist – they identify what the prompter wishes to have depicted, but the machine determines how those instructions are implemented in its output."
> "The USCO will consider content created using AI if a human author has crafted something beyond the machine's direct output. A digital artwork that was formed from a prompt, and then edited further using Photoshop, for example, is more likely to be accepted by the office. The initial image created using AI would not be copyrightable, but the final product produced by the artist might be."
Then where is the line? If I use a painting program that runs a Math.random() on a brush and I use this brush to draw something then is that also created by a non-human?
If I program a robot arm to draw a picture then is that picture also created by a non-human? What if I use a printer?
From what I gather, they just don't consider it a creation in a way that pertains to copyright laws. Like a rock on the beach.
It's a tool, not a conscious being with rights.
Let's be realistic here: there is no oracle indicating something was generated by AI. Most graphic artists have been using partially machine generated work from tools like Photoshop for a decade or more. They just don't mention it.
https://copyright.gov/docs/zarya-of-the-dawn.pdf is the letter that lays out why they don't think Midjourney is a paintbrush rather than a monkey.
I can't tell that this is a long-standing interpretation, rather than a recent ruling, conceivably a blip.
That's not the status quo. The biggest concern is that AI can sometimes generate code from the training set verbatim, in which case that code has the copyright of the original training set - which might be GPL or proprietary or whatever else.
Not just verbatim: AI-generated code is arguably a derivative work of code in the training set even if it doesn't generate verbatim copies of training data, in the same way that any code can be a derivative work even if it doesn't contain bitwise-identical lines.
If you take a piece of Python code, and translate it to C code, without a single line being copied verbatim, the resulting C code is still almost certainly a derivative work of the Python code.
Just out curiosity, what makes the code you write, more original than an LLMs who’s training set is way bigger than yours and likely will more variance in how to achieve same goal.
I’m not trying to be facetious but rather stoke the philosophical idea of the reality that humans aren’t as special and unique as we think we are.
Also, there's the possibility that all things creative are the human form of peacocks' tails, and if that's the case then the cost and difficulty matters more than the outcome, and any discussion based on capitalist incentives is fundamentally flawed.
https://kitsunesoftware.wordpress.com/2022/10/09/an-end-to-c...
(I wrote this 52 days before ChatGPT came out, the reference to GPT-3 is based on the stuff OpenAI released before the chat interface).
I don’t even think it’s factually true.
Along the way, you’ve experienced trauma which nudged you a direction, you’ve been inspired by other people which left an impression in your mind, you’ve had people teach you 1 skill and another person teach you a different skill, and you combined them, life experiences that changed your outlook etc.
It’s less about the specific teachers and more about every single person and event that has happened to you is your own ‘training set’. But sometimes we can definitely point to a specific event or person who influenced us to where we are at today and I like to think of that just like model weights in current LLMs.
Indeed in many ways LLM code is more original than mine, as I limit myself to calling only functions that exist and producing code that will compile, unlike LLMs which have no such limitations. /s
That would be like assuming every piece of code you’ve ever written compiled perfectly on every run. Do you hold yourself to the same standard? Or do you give yourself some leeway because you walk through it, talk through, think through it and then come to a solution?
Someone wrote the first murder mystery. It wasn't something we inherited from amoebas.
Someone wrote the first Regency romance. In fact, we know who it was - Georgette Heyer.
Someone invented calculus. We didn't get that passed down from the cave men via oral history or tribal knowledge.
We've dreamed of flying for millenia, but someone invented the first practical airplane - someone specific. Yes, they built on previous knowledge. But their step was still original. It had never been done before.
And so on, for idea after idea and original creation after original creation. They all originated at some time, with someone.
So at what point do we stop attributing the previous innovations that led to our current innovations? And why do we stop there? Would Mr. Caveman be the only special human to ever figure that out or could the argument be made that eventually someone else would have figured out how to make tools and therefore attribution is just pointless?
What I am trying to get at is, everything you do or create, is because of work the all of humanity has done. So pertaining to copyright, why should 1 single person claim an idea as solely theirs when that idea was not created in a vacuum.
I should also say, I am not discounting anyone’s work, but rather if the monetary reasons for creating became secondary, would the need for copyright even exist?
While I also think there is a good chance that what you are claiming (that any output is a derived work of the training set), this is definitely not settled, and will require someone going to court over it. I expect the proceedings over something like this to take a long time, and very likely reach to the supreme court. So, we won't know if you're right or not for quite a while.
A good argument against the idea that any output is a derived work of (all) data in the training set is exactly that verbatim copies of snippets from the Linux kernel can't be a derived work of snippets from GCC, even if GCC was also in the training set. So it is arguable that someone intending to claim copyright infringement for a piece of code generated by an LLM has to identify which particular work of theirs from the training set it infringes on, it can't be a blanket finding.
It won't be over when the first top-level court decides.
https://www.reuters.com/legal/ai-created-images-lose-us-copy...
so it may not be "public domain", but if you can't get copyright protection in the US I don't see a difference.
Barring a court case, we’re more in an uncertain place. I think it’s more likely for courts to hold that many model uses aren’t transformative enough to get a fair use exemption. But until a court says either way, it most definitely does NOT just default to public domain.
1. Does the output infringe an existing copyright?
2. Is the output copyrightable?
As far as the US Copyright Office is concerned. I do not believe other countries have weighed in yet.
Also, when your training data is based on copyleft source code, and in some cases is a direct ripoff of it, will that license be applied to it?
Ask an LLM how to do something you already know very well how to do properly, only then can you see its flaws through its mindless bravado.
> that I could write myself
I don't think your comment is very relevant, given how they specified it.
One person says "write code" in the context of chatGPT and they mean Common Lisp and chatGPT3.5
Someone else says "write code" and they mean React components and chatGPT4.
It is a mirror to how imprecise and intellectually lazy our own language and minds have become in the context of online discussion. Compressed thinking in order to fit in a twitter box along with years of being rewarded with attention the more muddled the thinking since the more people that disagree with the point the more responses it will generate.
It is like we have been running a giant language miscommunication RLHF model and the end result is an incredibly accurate system at creating miscommunication.
[1] https://ai.googleblog.com/2006/06/extra-extra-read-all-about...
GPT-produced code may have different failure statistics, and therefore the human heuristic may not work for GPT-produced code. It's too early to tell.
Unless I'm feeling particularly lazy or the code isn't important.
It's an interesting divergence in software, between those who manage complexity by adding more human-understandable abstraction, and those who manage it by just verifying the results, letting the complexity fly free. All the ML stuff is definitely taking big steps down the latter path.
What people perceive as a quality car, is often a lack of understanding. Some think they are driving a fine car, but someone who pays attention notices that the door gaps are not uniform, that the plastic on the inside will start to smell bad in summer, that you will get a lot more fatigue because of wind noise and tire noise, that the sound of the door being slammed shut is very hollow instead of a nice THUNK, that the automatic gear box often eats its gears, especially when driving slow, ...
There are many details normal people don't even notice about their cars, and then they claim it is a quality car. It is not, I am sorry, and it is the perfect analogy for this ChatGPT/Bard being a code generator.
If the user doesn't notice the door gaps, then it isn't a quality issue. If the user doesn't drive the car enough to wear out the suboptimal gearbox, then it isn't a quality issue.
Yes, this means that the definition of quality depends on the needs of the user.
Getting these things right takes extra effort. If the customer didn't care, it would not be done.
Certainly, not every layman will be able tell you in so many words WHY it's nice and everyone cares about different things in a product (while it's the designers job to care about all of them)
But that was not really the contention now, was it.
Google warns staff about chatbots
https://news.ycombinator.com/item?id=36341188 (107 comments)
> The policy isn't surprising, given the chocolate factory also advised users not to include sensitive information in their conversations with Bard
"Chocolate factory"??
There is nothing in the context of the article related to chocolate factories. AIs usually don't just put random words in an otherwise mostly fine article.
AI tends to have the opposite problem. The form is good, the words go well together, but it is nonsense when you look at the big picture.
Gooogle telling staff to not use Bard would be like telling them a decade ago “don’t use android”
bard.google.com states right on the front page: "Bard is an experiment and may give inaccurate or inappropriate responses."
Also, the warning is about _all_ LLMs.
Leave other LLMs to the side on this issue. There are a variety of different issue that arise when Google employees might use a non-Google LLM, so that prohibition is in a category with different reasons.
What makes you think so? Those LLMs aren't primarily coding tools.
> Leave other LLMs to the side on this issue
https://www.reuters.com/technology/google-one-ais-biggest-ba... indicates that the ban is about "all LLMs, period".
It's ElReg who, true to their style, decided to focus the discussion on Google employees using (or not using) Google's LLM. Even they, however, acknowledge that Google "also advised users not to include sensitive information in their conversations with Bard in an updated privacy notice."
As such, Google seems to apply the same standard to other users (including "countless /other/ coders") that it applies to its employees, no?
It doesn't have to be a primary use case for it to be one of the use cases Google is putting forward. In which case Google is still saying they're unwilling to use their own tools in the way put forward for others to do so.
"Hey world, Bard can generate code so well now, we have to restrain our people from using it!"
Something totally transparent can come out of Google after all.
Is a 30 year old company with over $1bn in revenue still a startup?