GitHub Copilot and the Methods of Rationality (2022)
freshpaint.io
freshpaint.io
If you’re getting inconsistent results, it’s most likely you’re doing experiments in a single project, and opening a “new file” expecting it to be a “clean slate” like a “new chat” in chat gpt.
…but codex doesn’t work that way.
If you find it “suddenly” figures out how to do something, particularly something you’ve done before, it’s most likely your solution is being fed in as part of the prompt, not that you’ve figured out some magical prompting keywords. Copilot just isn’t that good at code gen.
Personally I’ve found that codex is very predictable in terms of performance and consistency, but, broadly speaking, an order of magnitude less capable than chatgpt.
[1] - https://thakkarparth007.github.io/copilot-explorer/posts/cop...
It’s wild how much the narrative has changed around copilot since the release of ChatGPT. When copilot first came out there was so much breathless hemming and hawing over how it was so problematic for violating licenses. Turns out, it’s not really good enough to even be used in such a way that it could do that
Copy paste a function signature… Copyright issue… Ask for example Google about that. (I mean the company, not necessary the search engine.)
The class actions against MS & Friends regarding Copilot are still rolling:
that's a very different thing than copying a method signature here or there, and the suit was about the collection as a whole and would not have made it into a courtroom if it were just a few method signatures.
A nit, though: codex is the model itself, and doesn’t do anything to slurp up nearby files. That work is entirely done client-side, and is part of the “UI” of the tool. Code is heuristically scanned for relevant snippets using an 11-dimensional logistic regression model which is bundled into the copilot source code and executes locally, and those snippets are included automatically as preamble comments in the prompt. Thakkar Parth’s work showed this and you can further explore it by running a MITM proxy to capture the network activity and look at the prompts directly.
(Which, to be clear, is perfectly fine and useful, but it's weird that the author treats his process with so much reverence when he's not actually following the formal parts.)
By nexting, I think you know what I mean. What a reasonably intelligent junior would figure out you were doing, once you started doing it. Like if you have a bunch of cases in a match or if-tree, and you started doing the first one, he would know what you were up to and more or less get it right.
Anything more complex, I've occasionally been positively surprised, but mostly it's just plain wrong, though obviously wrong so it didn't throw me off.
I find that people have wildly different expectations about what LLMs will do for us, and this colours whether they are satisfied or not. For me, I'm very happy with copilot, it definitely saves me time, but I also wasn't expecting it to do much beyond guessing my next lines. I always thought I'd be the one to understand the grand structure and it would be the fast typing graduate, and that's how it turned out.
# Compute the difference between each element in an array.
# compute_difference([1, 2, 5, 10]) == [1, 3, 5])
Overall you shouldn't expect Copilot to read your mind. If your description is vague, you're going to get unreliable results. If you have a lot of context, clear descriptions, unit tests, existing code, etc, it has a better chance of guessing your intent.ChatGPT is far better at this sort of "describe what you want and it will code for you" process from TFA. It looks like the next version of Copilot with the chat interface will bring the latter workflow into the system, I'm looking forward to that, a GPT-4 based copilot.
My main complaint about Copilot is that it doesn't seem to feed Intellisense prompts into the prediction, or validate predictions against them. That will make it so much more accurate when they fix that.
Just breaking a problem down comment by comment and letting Copilot try to translate as you do so is super effective for getting oriented. It knows the syntax, and you can easily tell after the fact if it’s doing the right thing. But it’s helped me discover a ton of features that I would have missed and replicated by hand in the old days.
PEP 257 prescribes the imperative mood:
The docstring is a phrase ending in a period. It prescribes the function or method’s effect as a command (“Do this”, “Return that”), not as a description; e.g. don’t write “Returns the pathname …”.
GPT 3 is “fancy tab-complete”.
GPT 4 will correct your mistakes, write the test case to prevent your fallible human brain from repeating it, and then explain this in a detailed doc-comment so other humans can read it and gawk at the history of your errors.
I went through a process of re-writing my prompt several times to get correct output. In fairness, I think getting the correct result meant being more precise with the prompting language.
Overall, it was a bit spooky seeing it respond to my varying prompts; then noticing it's errors; then re-prompting to fix its misconceptions. In the end, though I could have written it myself faster, so I am not sure it is saving me time.
I am surprised when I see people saying its saving them massive amounts of time. Are there any good examples out there of it producing copious amounts of code, correctly, from prompting such that I can easily believe using CoPilot is a time-saver?
I saved a lot of time when I needed to write a whole bunch of relatively simple validators using a library I hadn't used before. I have no doubt I could have written the code myself, but using AI meant I could write a simple prompt like "Write a validator using library x to validate the following data" and then just cut'n'paste all the different data samples.
It was simple but boring, and a job I would only do once so learning the validation library's API wasn't worthwhile. I reckon I saved myself a day or two.
Intuitively in my mind, the way to use tools related to software development is to be as explicit as possible. LLMs are good at either translation, when the proper context is provided, and synthesis, when the context is not provided.
Once this is understood then the user realizes there are no “magic” words.
"""
The file named input in the current directory has a list of numbers in it, one per line. Each number
indicates the depth of the sea floor scanning outward from a submarine.
The first order of business is to figure out how quickly the depth increases, just so you know what
you're dealing with - you never know if the keys will get carried into deeper water by an ocean
current or a fish or something.
To do this, count the number of times a depth measurement increases from the previous measurement.
(There is no measurement before the first measurement.)
For example, suppose you had the following report:
199
200
208
210
200
207
240
269
260
263
In the example above, the changes are as follows:
199 (N/A - no previous measurement)
200 (increased)
208 (increased)
210 (increased)
200 (decreased)
207 (increased)
240 (increased)
269 (increased)
260 (decreased)
263 (increased)
In this example, there are 7 measurements that are larger than the previous measurement.
"""
--- CoPilot Output Synthesizing 9/10 solutions (Duplicates hidden)
=======
Suggestion 1
import sys
=======
Suggestion 2
def main():
"""
Main function
"""
with open("input", "r") as input_file:
previous = int(input_file.readline())
increases = 0
for line in input_file:
current = int(line)
if current > previous:
increases += 1
previous = current
print(increases)
=======
Suggestion 3
with open("input", "r") as f:
lines = f.readlines()
lines = [int(line.strip()) for line in lines]
count = 0
for i in range(1, len(lines)):
if lines[i] > lines[i-1]:
count += 1
print(count)
---Suggestion 1 probably completes into something reasonable eventually
Suggestion 2 works if you add a call to main()
Suggestion 3 works as-is
---
How would it even know the end user wanted python with that prompt below?
# Read an array from input.txt.
Am I missing some context, like it was called from vscode within an open project that already had some metadata or other context?As far as Copilot is concerned: Can you spell "copyright violation"? The T&C on GitHub do state that you grant them the right to "parse [your content] into a search index or otherwise analyze it on our servers". It's not at all clear that this grants them the right to reproduce parts of your content (without credit) using Copilot.
What about private repositories? "GitHub considers the contents of private repositories to be confidential to you." It would be interesting to see if one can get Copilot to produce code that is in a private repo.
It’s an even better experience with notebooks where the abstract vs repeat tension leans more toward the latter.
Here are three "homework-looking" functions that were generated by Copilot based on their signatures, with comments added by me:
# "Load" and "file" instead of more precise terms.
def load_numbers(file):
with open(file) as f:
# both .strip() and .readlines() are redundant and should be removed.
return [int(line.strip()) for line in f.readlines()]
# Verbose and imprecise naming.
def calculate_differences(numbers):
numbers = sorted(numbers) # Inefficient cloning instead of .sort().
differences = [numbers[0]] # Breaks on empty lists.
for i in range(1, len(numbers)):
differences.append(numbers[i] - numbers[i-1])
differences.append(3) # Where did 3 come from?
return differences
# Vague. Redundancy with param name. Sounds too much like English.
def count_increasing_numbers(numbers):
increasing = 0
for i in range(1, len(numbers)):
# Homework looking if/else.
if numbers[i] - numbers[i-1] == 1:
increasing += 1
else:
increasing = 0 # Why would it reset?
return increasing
And here are three similar functions from signatures where I tried to give better "vibes": # "parse" and "filename" are more precise, and out params indicate higher performance code.
def parse_numbers(filename, out_numbers):
# Explicit file-mode!
with open(filename, 'r') as f:
# Reading lines by iterating over the file object!
for line in f:
# Explicit error handling!
try:
# No redundant .strip()!
out_numbers.append(int(line))
except ValueError:
pass
# "get", "adjacent", and "diff" all show familiarity with vocabulary.
def get_adjacent_diff(numbers):
# Simple one liners! No off-by-1 errors!
return [numbers[i] - numbers[i-1] for i in range(1, len(numbers))]
# "Pairwise" is more technical, and avoiding name/arg redundancy. Less like English.
def count_pairwise_increasing(numbers):
# Another good one liner, making full use of list comprehension.
return sum([1 for i in range(1, len(numbers)) if numbers[i] - numbers[i-1] == 1])
Unfortunately this is not objective at all, and subject to changes on GitHub's whims.