Copilot Internals
thakkarparth007.github.io
thakkarparth007.github.io
https://github.com/search?q=GH_COPILOT_TOKEN&type=code
(My Copilot was broken and this was in the error output I was seeing, see: https://github.com/community/community/discussions/41878)
What I found there was some truly impressive reverse-engineering work by a single individual. I really like the "JOURNAL" daily-diary they kept of progress and random thoughts so you could see the progression day-by-day.
--------
One thing I found interesting: The author says that it queries only the 20 recently most opened files of the same language.
But in an AMA, I asked about how much "context" Copilot has available, and one of the devs says it can, for example, read header files that pair with C/C++ files that are open in separate tabs:
https://github.com/orgs/community/discussions/29932#discussi...
> "I assume Copilot uses the contents of the current project (IE, all files) as contextual information to offer suggestions. Is this the case?"
> "Yes, copilot looks at what we call "related tabs", for example .c files are often paired with .h files, so copilot considers the contents of other tabs open in the editor to try to feed a more complete prompt to the model."I'd love to know if you guys have any specific questions about copilot's internals that I can try to answer by staring at the code or if you have any feedback for the tool/post!
> Then, most recently accessed 20 files of the same language are queried from VSCode.
are probably not available by default in nvim?
---
Also, do you think there's any chance you get in trouble for reverse engineering copilot?
Looks like the 20 file logic is present in neovim as well. I went through its code (https://github.com/github/copilot.vim/blob/release/copilot/d...) after beautifying it (https://codebeautify.org/jsviewer), and found it present.
I couldn't exactly trace it to a specific neovim event but I'm guessing it corresponds to buffer-update-events (https://neovim.io/doc/user/api.html#api-buffer-updates) or something like that.
Re: getting in trouble
I surely hope not :P. I mean, the code is basically public (available on every user's computer).
I played with ChatGPT and asked it interview questions, and I thought it was a pretty interesting exercise to find its mistakes and get it to fix them. Good tool for training interviewers, perhaps.
However, I do believe there could be a meta model that can query code and libraries.
Does anyone believe that?
edit: I'm surprised to see that (so far) 3 replies actually agree with the statement. Is there a video that you'd recommend that shows realistic usage and gain from copilot? Maybe a livestream or something.
In terms of design, I had a long conversation with ChatGPT the other day about designing a database, including optimizations that could be made given certain requirements and constraints, etc. It was a big productivity boost, like rubber ducking on steroids.
I could not think how it would have helped me, but maybe I m limited in my imagination or don’t know how to ask.
At one point I told it to name the algorithm we had been discussing something like "OptSwim" and we just kept iterating on the idea.
It routinely invents arguments, functions or concepts which don't exist in reality or don't apply to the current context, but look like they could, so you are even more likely to get caught by this.
It's just taking the "I wish they'd thought of my use case when designing that API" on the next level by simply pretending in a very sincere and convincing way that your wish came true, then writing a usually-pretty-correct program around that assumption that would actually work _if that wish had come true_ - but unfortunately that API doesn't really accept this convenient parameter, so...it's not that easy in reality.
However, for anything that requires me to think, it's 5% at best.
Don't take up the 50% figure as anything serious, I think it's just a way to state "if it is a such a meaningful boost in productivity".
Which it is, for a lot of tasks, because the vast majority of programming jobs are boring stuff outside of the HN bubble.
It's amazing how much of the world economy runs on csv uploaded to ftp servers.
It does mundane work exceptionally well
But yes, mundane work it is best at. Some things I have found it made particularly easy:
- scraping websites
- file i/o
- "mirroring" things (I write a bunch of code for doing something on x axis, it automatically replicates it for y and z etc with the right adjustments, or cardinal directions, or arrow keys, etc etc etc)
It wasn’t something I actually needed help on, though. When I tried to go further with it and complete more of the task, it got stuck in a loop of just suggesting more and more comments but never offering more code, and then it mysteriously stopped responding at all.
This is the best experience with it I’ve had so far.
At first I found it very useful to ask it to parse the input. Much faster than looking up three separate docs to piece together what I had in mind.
But then I asked it to parse a more complex input and it just kept failing badly even when I gave it sample inputs and outputs.
I’d say it definitely offers some productivity gains and is worth trying.
This year I recorded most of my days and uploaded them to youtube. So if you want to get a realistic view, take a look here: https://www.youtube.com/channel/UCOqPGQCzgieAOL6iOJjj8hg.
The earlier days you can see it speeds you up a lot. The later days (such as today) you still want to wrap your own head around difficult computer science concepts so it is kind of useless.
Let me know if you have any questions!
Also, even with "Chinchilla laws", you still gain performance in a larger model, you just need a lot more data (if just as noisy) to reach the same level of convergence, but a larger model will have already partially converged to a superior model with the same amount data.
Not true. See figure 2: https://arxiv.org/pdf/2203.15556.pdf#page=5
The loss decreases with greater model size at the same compute budget (i.e. stopping sooner regarding training data). Also some rehearsal/multi-epoch training improves the forgetting rate (thereby improving performance substantially), which hasn't been taken into account by Chinchilla et al. because they train <1 epoch.
Their text about Figure 3 confirms what I'm saying: "We find a clear valley in loss, meaning that for a given FLOP budget there is an optimal model to train"
But Copilot is getting tons of live editing data from its users too, and soon should be able to construct a nice dataset of edits. There's no way they aren't already doing that.
The live data is gonna be useful though ya. Is Copilot allowed to use it though under ToS?
> User Engagement Data When you use GitHub Copilot it will collect usage information about events generated when interacting with the IDE or editor. These events include user edit actions like completions accepted and dismissed, and error and general usage data to identify metrics like latency and features engagement. This information may include personal data, such as pseudonymous identifiers.
> Code Snippets Data Depending on your preferred telemetry settings, GitHub Copilot may also collect and retain the following, collectively referred to as “code snippets”: source code that you are editing, related files and other files open in the same IDE or editor, URLs of repositories and files path.
That's useful but I edit code a lot. And if I have 10 similar lines and made one edit, it'd be very convenient for Copilot to suggest edit following line or even lines.
https://crfm-models.stanford.edu/static/help.html
The next stronger Codex model is called code-davinci-001 which appears to be a fine-tuned version of the GPT-3 Davinci model which is known to have 175B parameters. The model naming is alphabetical in the order of the model size:
https://blog.eleuther.ai/gpt3-model-sizes/
See also A.2 here: https://arxiv.org/pdf/2204.00498.pdf#page=6
[0] https://beta.openai.com/docs/model-index-for-researchers
The plugin for IntelliJ (PyCharm etc), is this written in Java? Reverse compiling this might give some additional insights.