Why is that surprising to you? CoPolit doesn't actually know how to code, it just generates symbols that match learned patterns.
Sometimes these generated symbols don't represent valid code and since CoPilot doesn't actually perform filtering based on syntax checks or JIT, these results end up as suggestions.
This is actually a point where future versions could greatly improve the usefulness, e.g. use the compiler infrastructure to verify and filter generated results.
This includes auto-formatting and even result scoring by code metrics (conciseness, complexity, ...). Plenty of room for improvement even without touching the underlying model.
Looking at it the other way around, had they chose _perfect_ examples there’d be similar criticism about “cherry picking”.
It was in a file called "find_files.sh", and the command it generated was something like:
> find .... -exec ... {}\;
The problem is that there's no whitespace between the {} and \; and this is rejected by both GNU and BSD find.
It might've just been a mistake that someone made while putting the landing page together. But if it was generated, then how?? I can't imagine there are many (if any) instances of "{}\;" in the training data, since it's not a valid invocation of the find command...
#!/bin/bash
# List all python source files which are more than 1KB and contain the word "copilot".
find . \
-name "\*.py" \
-size +1000 \
-exec grep -n copilot {}\;
[1] https://web.archive.org/web/20210701061847if_/https://copilo...- This lists the lines containing copilot instead of the files. You want `grep -l` to list the files.
- `--size +1000` finds files with over 1000 blocks. `--size +1000c` gets you files over 1KB.
- `-name "*.py"` only finds files literally named `.py`. To find files that match the glob `.py`, you would single quote it like `-name '*.py'`
Idea: make Copilot test-driven. You provide a function invocation and an example output, and it searches for code with similar test cases, then mutate the code until the test case passes within its own runtime (it may need to mock certain things to make tests pass within its own runtime).
The examples all look cherry picked for presentation; but the ones who had access to Copilot early didn't even stress test it, yet it still generated garbage. And it took 1 in 10 tries to get it right. [1]
Of course, we'll hide behind the 'Technical Preview' argument but since the core of this 'AI' (GPT-3) is a black box system, we'll probably never know why it is generating the code it is generating.