Gitlab’s AI-assisted code suggestions
about.gitlab.com
about.gitlab.com
I actually feel like the suggestions seemed to get worse during my month of using it for some reason.
Towards the end, it started suggesting these large blocks of code (another issue I had with the interface as well) which were very not relevant to what I was attempting to write.
All in all, I’m very underwhelmed from my first experience with “AI enhanced” coding.
This is all it takes for me to get value out of my subscription. It's saving me time by anticipating my next test / block of code based on what I've previously written, and that saved time really adds up.
Sure it's not "Jesus take the wheel" level AI yet, but time is money, and so is my sanity.
It really needs to be something where you can rattle off a list of requirements and have it build the code for you. The code context is necessary, but not sufficient.
Also, if you enhance your available information with AI, why not just write no message at all (or a short one as always) and use the AI when _looking_ at the commits/PRs? The AI will certainly be better at this in a later stage because they themselves will get better and they will have more context due to being able to look at the commits that followed the PR/commit.
There is no point in generating the commit message when committing.
When auto-generated messages will become a commodity, there will be issues with it. For example, if the message doesn't fit with the actual commit contents (semantically or syntactically), I have to think hard whether this is an AI-message-generator-bug, or the author missed something or wether I am missing something. This is not cool and makes reading commit messages harder.
Learning how to get the best results of them takes a great deal of experimentation.
I'm increasingly using my LLM CLI tool for quick lookups, and sometimes to make changes to existing code too - see https://github.com/simonw/symbex/releases/tag/1.0
Do you not use the direct vscode integration? I don't think I'd use it otherwise. It's been hugely convenient with.
I get great results from Copilot, but that's because I do a bunch of things to help it work for me.
One example: I'll often paste in a big chunk of text - the class definition for an ORM model for example - then use Copilot to write code that uses that class, then delete the code I pasted in again later.
Or I'll add a few lines of comments describing what I'm about to do, and let Copilot write that code for me.
Did you try the trick where you paste in a bunch of code and then start writing tests for it and Copilot spits out the test code for you?
A few more notes about how I've used Copilot here:
- https://til.simonwillison.net/gpt3/writing-test-with-copilot
- https://til.simonwillison.net/gpt3/reformatting-text-with-co...
I've also used Copilot to take educated guesses at things - like an inline mini-ChatGPT - with good results:
- https://til.simonwillison.net/gpt3/guessing-amazon-urls
As with so many other AI tools, Copilot is desperately lacking detailed documentation. It's not at all obvious how to get the most out of it.
Even something as simple as adding a doc block for a method can boost the quality quite a bit.
I could live without it, but I'd also keep paying if the price increased.
I wonder if this quest for the perfect prompt is dumbing me down, though.
There are a lot of tricks to optimizing output from Copilot ("Copilot-Driven Development"), I wrote a little about what I've discovered here:
https://gavinray97.github.io/blog/a-day-without-a-copilot#co...
I don't care that much if I sound pretentious or too harsh, but how is this not common sense? Maybe the general standards in the industry are way lower than they used to be. This looks like when one of the consultants hired by our company and paid 10k+ per month for a ISO 27001 certification said: you need to make sure your password fields show a * instead of the actual character.
After using copilot for 2 months, I found it somehow useful for very generic repetitive cases, but for complex specific stuff it was very bad. I'd rather spend an hour writing code and understanding the problem I'm trying to solve than spending an hour understanding and adapting some generic code that partially does what it needs to do. And if there are people out there that fully trust it and then deploy code that they don't understand and can't explain themselves, then we're in for a lot of fun in the future.
I was hooking IAudioClient COM class to capture and silence arbitrary app's audio, and as soon as I wrote signatures for its members Copilot was able to generate skeletal stub implementation with logging as well as hooking code totaling about 150 lines of Rust.
In addition, Copilot benefits from a large training set, which I guess is best for Python and JavaScript as two very popular languages. Strictly typed languages seem like they are generally less well represented, at least among public GitHub projects.
If I write a comment describing what I’m doing, it will generate multiple lines of correct code
The first two points in @gavinray’s methodology being unfortunately critical.
I think it's biggest benefit isn't for me, though. Copilot is great at churning out predictable or repetitive stuff, much less in writing a abstractions over or around that repetitive or predictable stuff. E.g. it's great at writing yet another CreditcardPayment ActiveRecord model for Rails. But less so for an obscure business domain model used outside of a framework. And great at writing fifteen getters and setters, less for an abstract base class, DSL or macro that introduces all these getters and setters for me.
It's also bad at certain rewrites. A codebase was slowly rewritten from OldStyle (and old libs) to NewStyle. It stubbornly kept suggesting OldStyle snippets.
And last, I find Copilot has a far higher ROI in some languages and frameworks than others. E.g. dynamic languages (like Ruby) have very poor LSP and intellisense support compared to typed and static (like Rust). So the added benefit differs a lot based on where it's used, is my experience.
I guess esp. the latter is why I too am underwhelmed. But also why I'll keep using it, this time, for when I do the inevitable Rails, JavaScript or Panda's gigs. But less so for my weird eventsourced, hexagonal, Rust project.
Also its inability to match parens and braces/brackets is legendary. So annoying.
But! It is extremely helpful for filling in error messages. I don't think I wrote a `throw new ` line fully since I started using it. The quality and information density of my error messages has increased considerably.
Not sure if it is worth keeping for just that but it is a nice benefit.
Only reason I still have it is because the beta claims to have GPT-4 integration and that could prove to be a very powerful feature. But I still haven;t gotten an invite to install.
But honestly for me it reduces the attractiveness of gitlab. Just adding half working features does not increase the value for me.
Please finish and polish the existing Features before adding hype feature 93939. Maybe add a policy that your open feature and big count should be less than 20K or something.
Examples: Security scanning has a god awful UI and UX.
The whole project management stuff is barely usable
Pipelines have loads of weird half features and missing QOL. E G. Adding saeif support instead of just your own warnings format.
I really like Gitlab, I think the CI experience is a lot better than, for instance, Github Actions, but I also don't ever see me using a good chunk of the stuff that's been shoveled on in the last few years and I don't really know anyone that does.
Gitlab did not support poetry for ages there where bugs with python galore and for dotnet they do not speak the same language - e.g. issues in dotnet are reported as saeif, but gl has it own weird format.
And yes, you are totally right with the bug tracker. It's often funny how long the ticket is just because it was retagged sooo many times.
The training data is documented in https://docs.gitlab.com/ee/user/project/repository/code_sugg...
AI Transparency is important, all available AI features provide documentation for training data, and are built with privacy first.
The GitLab Duo announcement adds more feature details and plans. https://about.gitlab.com/blog/2023/06/22/meet-gitlab-duo-the...
The AI/ML blog series provides insights on experiments, and features being built. https://about.gitlab.com/blog/2023/04/24/ai-ml-in-devsecops-...
Would it be possible to get a complete list of sources and licenses?
There's not actually anything in GPL or any other major FOSS license that prohibits using it to train an AI. The controversy stems from Microsoft refusing to follow the terms of those licenses. If they could just fulfill their obligation to propagate copyright statements and license text there would be no controversy over copilot.
Thus stating that it was trained only on "permissive" licenses doesn't actually answer the question without defining what a permissive license is.
Most of the techniques I used, I've invented myself (or learnt from the standard documentation). When I use a technique that I haven't invented myself, I look up where it came from. Half the time, my version is actually radically different (and my attribution is mistaken); the other half, I've remembered an inferior version, so I steal the better version and then attribute it appropriately.
That's one way it's different. There are others. Really, though, we should be asking the question “in what way is this the same as humans learning?”, expecting answers that would convince an education specialist.
An LLM's working memory is just its context window, but the LLM also has embedded data in its parameters, which are set during training to minimize loss. This is effectively a memory of the training data, just as much as a digital photo of my face is a memory of the photons reflecting from my face.
AI is literally 1-1 equivalent to compression. If you don't believe me, you should check out this demo of GPT-2 as a (for awhile SOTA) text compressor.
Oh its down now, but here's the HN thread to prove this exists: https://news.ycombinator.com/item?id=23618465
This is a very strong statement considering most code is just rearranging existing patterns. Unless you are doing cutting edge academic research, I'm very skeptical of your claim.
> (or learnt from the standard documentation)
This means my Python code is mostly Pythonic. (The rest is kinda idiosyncratic, but I don't often get complaints.) But also:
> considering most code is just rearranging existing patterns.
I am an outspoken critic of "design patterns". They have their place, if you're working with legacy tooling like C++, Java or Rust, but if most of your work is rearranging existing patterns, you have long outgrown your tooling and you need a better programming language. (Or you're copy-paste programming and need to learn your tooling first.)
I am doing academic research, but that's besides the point. So far, I've learned that it's rather hard to be cutting-edge if you don't look at other people's work: you end up re-inventing all the wheels, and any insight you may have brought is lost in the noise.
Imagine a prompt "Photo of person, Shutterstock ID 132456, with blue eyes instead of brown eyes, watermark removed"
If the prompt returns Shutterstock photo #123456 without the watermark (and with the different color eyes) but otherwise a near identical photo, I think most people would agree the output shouldn't be free to use without buying the original photo license from shutterstock.
To a certain extent, we're betting that these models won't accept or reply to prompts that are that specific (e.g. referencing a specific image for sale on shutterstock by its ID number). Or even just providing the photo in the prompt and asking the model to remove the watermark and upscale the image to a higher resolution.
I'm fearful that LLM's will become (or already are?) an easy copyright bypass tool that can be abused, in the example above, to put companies like shutterstock out of business.
This is the sort of problem regulation might help with.
I haven't read OpenAI's TOS, but I'm curious who owns the output of the model and whether OpenAI is transferring copyright/licensing liability onto the user or if OpenAI is representing that output from the model is 100% free to be used in any way the user wants.
On the other hand, if a developer using this tool then goes and tells it "please write me a C library in the style of GNU libc," then yeah, that is skirting a fine line. But just don't do that.
// fast inverse square root
https://news.ycombinator.com/item?id=27710287Why not songs, software, entire books and tv shows?
Complete the function for an sht21 driver [code omitted]
What it returned was the copyrighted function, comments and all: static inline int sht21_rh_ticks_to_per_cent_mille(int ticks)
{
ticks &= ~0x0003; /* clear status bits /
/* Formula RH = -6 + 125 * SRH / 2^16 from data sheet 6.1,
* optimized for integer fixed point (3 digits) arithmetic
*/
return ((15625 * ticks) >> 13) - 6000;
}
Notice how it even included a comment referencing a specific datasheet!This driver is hardly "famous" or even "notable", because those aren't things LLMs understand. The prompt simply contains enough context to be distinctive and the sht21.c is an old, stable driver in each of the many kernel trees included in its training set.
Regurgitation isn't a particularly rare thing with LLMs, most cases just aren't this obvious.
[1] https://github.com/torvalds/linux/blob/c6b0271053e7a5ae57511...
> Google Vertex AI Codey APIs are not trained on private non-public GitLab customer or user data.
It prevents bad actors like Apple from ripping off people's philanthropic labor, but it also prevents me from ripping off people's labor. It also focuses effort onto the FOSS project.
I like PyQts solution of having GPL or buy a commercial license.
I suppose I still like MIT/Apache style the best. Even if someone rips them off, we didn't lose progress.
Compile already working code, slap my logo on it, spend millions of dollars marketing it with young good looking adults subliminally letting you know that you aren't cool unless you give me money.
But I also don't have the ethics to do this. You'd need a real psycho to do this...
Make your logo a trademark so that others can't use it. It doesn't violate GPL because trademarks are not copyright, and make sure your marketing campaign drives home the point that you are the real deal and all others are ripoffs (including the original).
I mean, people manage to sell bottled water to people who have perfectly good and 1000 times cheaper tap water.
If a human did what these language models are doing (output derivative works with the copyright and license stripped), it would be a license violation. When humans want to create a new implementation with clean IP, they have one team study the IP-encumbered code and write a spec, then a different team writes a new implementation according to the spec. LM developers could have similar practices, with separately-trained components that create an auditable intermediate representation and independently create new code based on that representation. The tech isn't up to that task and the LM authors think they're going to get away with laundering what would be plagiarism if a human did it.
I'm curious how easy or difficult it is to get GPT to spit out content (code or text) that could be considered obvious infringement.
Tempted to give it half of some closed-source or restrictive licensed code to see if it auto-completes the other half in a manner that is obviously recreating the original work.
Edit: it wasn't ChatGPT but Copilot see https://twitter.com/mitsuhiko/status/1410886329924194309
The same applies to GPT. It could reproduce Bohemian Rhapsody lyrics in the course of answering questions and there’s no automatic breach of copyright that’s taking place. It’s okay for GPT to know how a well known song goes.
If copilot ‘knows how some code goes’ and is able to complete it, how is that any different?
The weights are an intermediate representation that contains nothing resembling the original code.
https://sites.google.com/view/stablediffusion-with-brain/
I think brain to neural net alignment is justified by the fact that both are the result of the same language evolutionary process. We're not all that different from AIs, we just have better tools and environments, and evolutionary adaptation for some tasks.
Language is an evolutionary system, ideas are self replicators, they evolve parallel to humans. We depend on the accumulation of ideas, starting from scratch would be hard even for humans. A human alone with no language resources of any kind would be worse than a primitive.
The real source of intelligence is the language data from which both humans and AIs learn, model architecture is not very important. Two different people, with different neural wiring in the brain, or two different models, like GPT and T5 can learn the same task given the training set. What matters is the training data. It should be credited with the skills we and AIs obtain. Most of us live our whole lives at this level and never come up with an original idea, we're applying language to tasks like GPT.
At the same time it does not replace human developers in any application, it might take a long time until we can go on vacation and let AI solve our Jira tickets. Remember the Self Driving task has been under intense research for more than a decade now, and it's still far from L5.
It's a trend that holds in all fields. AI is a tool that stumbles without a human to wield it, it does not replace humans at all. But with each new capability it invites us to launch new products and create jobs. Human empowerment without human replacement is what we want, right?
You can't just take copyrighted code, base 64 it, sent it to someone, have them decode it, and claim there was no copyright violation.
From my (admittedly vague) understanding copyright law cares about the lineage of data, and I don't see how any reasonable interpretation could consider that the lineage doesn't pass through models.
IANAL
What if we train the model on paraphrases of the copyrighted code? The model can't reproduce exactly what it has not seen.
Also consider the size ratio - 1TB of code+text ends up into 1GB of model weights. There is no space to "memorize" the training set, it can only learn basic principles and how to combine them to generate code on demand.
The copyright law in principle should only protect expression, not ideas. As long as the model learns the underlying principles without copying the superficial form, it should be ok. That's my 2c
So is the ELF.
Maybe at a FAANG or some other MegaCorp, but most companies around barely have a single dev team at all, or if they're larger barely have one per project.
... and then execute copyrighted code -> trace resulting values -> tests for new code.
AI could do clean room reimplementation of any code to beef up the training set. It can also make sure the new code is different from the old code at ngram-level, so even by chance it should not look the same.
Would that hold up in court? Is it copyright laundering?
Potentially for all of the inputs at once.
What language models could do easily is to obfuscate better so the license violation is harder to prove. That's behavior laundering -- no amount of human obfuscation (e.g., synonym substitution, renaming variables, swapping out control structures) can turn a plagiarized work into one that isn't. If we (via regulators and courts) let the Altmans of the world pull their stunt, they're going to end up with a government-protected monopoly on plagiarism-laundering.
But IMO there are plenty of other places to add real value across the GitLab product with AI/ML features.
Here, it just looks like they saw GitHub do something and felt a need to copy it. But two years late, and worse.
As a longtime GitLab user (and onetime contributor!), I'm a bit disappointed they're spending so much time on this. I think they're just too far behind.
> But IMO there are plenty of other places to add real value across the GitLab product with AI/ML features.
True, and after starting with ML experiments, the product and engineering teams have been working on new features for entire DevOps lifecycle. All AI workflows on the DevSecOps platforms are described in the GitLab Duo announcement blog post https://about.gitlab.com/blog/2023/06/22/meet-gitlab-duo-the... and website https://about.gitlab.com/gitlab-duo/
I'll share a few highlights that I am personally excited about
- Explain and help fix security vulnerabilities. From my personal experience, I often find CVEs hard to read, especially when I am not the author of the code to fix. Getting help from AI can reduce entry barriers and make development for efficient. Security is everyone's responsibility these days. This follows the AI assisted feature to explain code in general. "What does this magic loop with memcpy do?" might not stay magic anymore, easing the path to code refactoring, improving performance, and reduce the resource usage footprint.
- Summarize issue comments. Feature proposals or bug analysis can have long comment threads that require reading time. AI will help get the gist and better contribute to what has been discussed.
- Summarize MR changes, to avoid reading long change diffs. This helps with faster (code) review cycles. I tested it this week with an MR for our handbook in https://gitlab.com/gitlab-com/www-gitlab-com/-/merge_request...
I'd also like to see AI helping fix CI/CD pipelines fast. Proposal in https://gitlab.com/gitlab-org/gitlab/-/issues/386863 I shared some thoughts in a new talk "Observability for Efficient DevSecOps Pipelines", slides in https://go.gitlab.com/VDAvMw (GitLab blog post coming soon, https://gitlab.com/gitlab-com/www-gitlab-com/-/issues/34296)
Additionally, I learned some new ideas at Cloudland last week, regarding product owner requirements list verification, and end-to-end test automation with AI. Need to create feature proposals :-)
> As a longtime GitLab user (and onetime contributor!),
Thanks for contributing. I'd like to invite you to share your ideas about AI features across the platform :)
When you look at the DevOps lifecycle (image in https://about.gitlab.com/gitlab-duo/) from plan/manage to create, verify, secure, package, release, deploy, monitor, govern - where do you see yourself, and where do you spend the most time in?
Second question: Which process feels the most inefficient? After identifying answers to the questions, please check the AI features https://docs.gitlab.com/ee/user/ai_features.html and/or open new feature proposals for GitLab https://gitlab.com/gitlab-org/gitlab/-/issues/new?issuable_t... You can tag @dnsmichi so I can engage with your ideas. Thanks!
I think it's more an internal itch. For a persistent money loser, the stock in holding there pretty well.
Showcasing innacuracies in AI responses must be a Google requirement.
I can't check how it looked 15 hours ago.
Makes me wonder if it was manually corrected and is thus a fake example now.
Our web team is working to resolve this issue here: https://gitlab.com/gitlab-com/marketing/digital-experience/b...
Way better than the AWS code whisperer or whatever it's called but still not worth switching. Especially since I trust GH way more than some random company (GH already has access to my code anyway)
But using it as a better autocomplete is actually quite nice, for example when I have some more repetitive code to write. And a line-by-line AI completer will at least give me the time to read it piecemeal, which I find much easier. Just like I understood way better when my math teacher demonstrated how to solve a problem compared to reading it from the textbook.
It's getting really annoying how many sites force you into a trial just to find out how much it'll cost when it ends.
EDIT: Is this even positioned to compete with Copilot? What editors are there plugins for? There is surprisingly little information on the site.
We offer experimental support for additional editors: https://docs.gitlab.com/ee/user/project/repository/code_sugg...
When GA, it will be included in our $9 per user per month AI add-on.
Signing up for the free trial funnels me to the trial of Gitlab Ultimate. So assuming that you need an Ultimate subscription to use it after the trial, that's the price. Pricing is here https://about.gitlab.com/pricing/
In contrast, Copilot is $10 per month https://github.com/features/copilot#pricing
It's fine to integrate Off-the-shelf solutions, but you don't have to implement every tech-hype going viral. I fear AI-assistant features are pushing more prescient features further down the backlog.
[1] https://docs.gitlab.com/ee/user/project/repository/code_sugg...
Could you elaborate on that? LLMs seem perfectly suited for language tasks like translation. They don't seem particularly expensive either, especially compared to hiring a person.
For a maybe more obvious example, say that LLMs ever got good enough to do arbitrary precision arithmetic on numbers up to hundreds of digits. Would that be a good use of one when calculators can already do this and are far cheaper to produce? I guess it makes no difference from a free-tier consumer's perspective, but it's still more expensive even if you aren't personally paying the expense.
Your analogy would be like saying why use a computer to multiply numbers if you can calculate them using calculator, which is much cheaper. Sure, but if you already have a computer, no need to use a dedicated calculator as wel.
[0] https://nlp.seas.harvard.edu/annotated-transformer/#results
https://www.reddit.com/r/Korean/comments/13lkh6c/gpt4_is_far...
https://github.com/ogkalu2/Human-parity-on-machine-translati...
[0] "Attention Is All You Need" https://arxiv.org/pdf/1706.03762.pdf
presumably to be used later to lookup words (crap approach, but it's an LLM, what do you expect)
but it didn't bother to write any code, it's just data
(plus it even managed to screw up the dictionary with double braces)
I have a free gitlab account but can't proceed without entering business information as it seems to be tied to the gitlab premium feature set. Am I doing something wrong ?
The training data is publicly documented, and AI features are built with privacy first.
For all URLs please check my comment in https://news.ycombinator.com/item?id=36526159
AI for self-managed instances is also focussing on privacy, while bringing more efficiency to teams. https://about.gitlab.com/blog/2023/06/15/self-managed-suppor...
https://docs.gitlab.com/ee/user/project/repository/code_sugg...