>> gpt-4 is at least as good as i am at code reviews, so i don't think this solves the problem this post is about
These two comments don't seem consistent.
Honestly, "i've found mistakes in its code reviews" vs "it's absolutely superhuman at is avoiding red flags" is possibly not self-consistent. But I think you mean glaring mistakes?
Which if I understand you correctly, than I'm not sure how you get the first conclusion unless you are saying you're mediocre (which is fine). But really it just makes it seem like you and the machine complement one another, rather than compete, which makes me not understand the original comment in context.
But I mostly agree with Copenjin. Interviews are less about on the spot skill checks so much as learning how someone thinks and problem solves. Honestly, you could play a boardgame/videogame/cardgame with them and it would be an effective interview (asking them relevant questions during the game, and maybe even better if it's a relatively unknown game so they can zero shot it). The reason for this is that in real world work you are more concerned with how someone adapts to changing environments and thinks through situations. To see when they'll ask for help, what they might get stuck on, and how they strategize.
Your employees will always be gaining new skills. And honestly, it is easier to take a lower skilled person who's more adaptable and driven and turn them into a great employee than it is to take someone who's got skills but will stagnate. But ymmv depending on the job and requirements. Sometimes you just need to fill a seat.
by 'red flags' i inferred copenjin to be referring to things like getting aggressive or defensive, rather than making dumb mistakes, but i could be wrong about that. i guess there are also some mistakes that are so dumb that they'd be a red flag, and i have to admit that gpt-4 is somewhat subhuman at avoiding those
if you play a board game with someone you can assess their general intelligence and capacity for logic. all else being equal, having more general intelligence and knowing how to think logically do make you a better programmer. (if you just want to assess general intelligence, a much faster pair of tests would be reverse digit span and reaction time.) but those are far from the only things that matter, they're not enough to be a great programmer, and they're not even among the most important factors. other important factors in programming include things like knowing how to program, knowing how to listen, and being willing to ask for help (and accept it), which a board game generally will not test
I agree with you. Which is why I say that they complement. But your reply to Baron implies that what they were suggesting wasn't a solution. I agree with the sentiment of the post to do things in person. But what I take from Baron is that it is much harder to fake the process with GPT because the actual part of the code review isn't so much about finding the bugs, it is you watching someone perform the code review (presumably through screen sharing and a video chat). You could have this completely virtual, but I think you're right to imply that the same task could be then optimized.
But at the end of the day, I think the underlying issue is that we're testing the wrong things. If GPT can do sufficient, then what do we need the human for? Well... the actual coding, logic, and nuance. So we need to really see how a human performs in those domains. Your interviewing process should adapt with the times. It is like having a calculus exam where you test someone and ban calculators but also include a lot of rote, mundane, and arduous arithmetic calculations. That isn't testing the material that the course is on and isn't making anyone a better mathematician, because any mathematician in the wild will still use a calculator (mathematicians and physicists are often far more concerned with symbolic manipulation than numerals).
> other important factors in programming include things like knowing how to program, knowing how to listen, and being willing to ask for help (and accept it), which a board game generally will not test
I agree the game won't help with the first part. But I thought I didn't need to explicitly state that you should also require a resume and ask for a github if they have one. But I did explicitly say you should ask relevant questions. And I'm not entirely convinced on the latter, which are difficult skills to check for under any setting. There's a large set of collaborative games in which do require working and thinking as a team. I was really just throwing a game out there as a joke, being more a stand-in for an arbitrary setting.
At the end of the day, interviewing is a hard process and there are no clear cut solutions. But hey, we had this conversation literally a week ago: https://news.ycombinator.com/item?id=40291828
Are you bad at code reviews? Is the code you're reviewing fairly standard?
GPT misses nuance. It can't reason while you still can. It certainly can do certain tasks better than you but certainly humans can do better at other tasks (specifically in the creativity side, logic, and when it comes to nuanced thinking). But if you're always focused on being quick (quantity over quality) then yeah, I think GPT could replace you. Otherwise, I don't know how anyone comes to this conclusion.
I've seen it notice logic errors in proprietary non-standard code that a human missed. It may not be able to literally "reason" through your code but it can follow the logic pretty well. I've even been able to have it pretty accurately comment the "intent" (vs. function) or spaghetti code with some reasonable accuracy.
gpt-4 makes stupid logic errors a lot, and i agree that sometimes it's bad at nuance, inappropriately applying heuristics in a context where they're inapplicable. it's much more creative than people are, though, and people also have those same flaws
"Can you scroll back a bit? Now, can you show me the docstrings for Foo.detachBar and Bar.attachFoo?"
I quite liked it, though I wasn't a fan of how bad the PR is itself, since I ended up having way too many comments, some of them being massive "wtf is this and why would you ever do it this way?" Types of things that IRL I would've refused to review without a proper rewrite.