Detecting potential cheaters in Advent of Code Leaderboards
adamfallon.com
adamfallon.com
I’m not sure that just fast Part2 is that strong a signal. (I’m also not even sure that I think GPT is cheating any more than having pre-written frameworks, ready Dykstra, A*, and min-max solvers, automated input downloaders, scripted submission functions, etc. I don’t think any of those are cheating. GPT is more a grey area, I guess.)
Is it basically the same kinds of CS exam puzzles each year?
In that sense, it has a definite pattern, but I wouldn’t say I experience it as formulaic.
Every year I curse a lot, multiple times.
It is basically a CS class.
But the nice thing is that most puzzles have multiple solutions, it's always educational to read how other people did things!
No? isn't that just like Google "i'm feeling lucky" but without even bothering to look up who said what and why.
GPT strips things of context, and this makes it difficult to know how much to trust it, and value it in comparison to alternatives.
I mean if you don't mind quake3 code turning up in your work, and don't mind even knowing that's what it is and where it came from then i guess that level of don't give a crap matches what you are doing - seriously if it's just a hack i have nothing against this, but it doesn't feel right for any long living code.
Can I claim this guy is a fraud for using python instead of typing opcodes into a hex editor? He doesn’t even know what registers hold his data!
These tools seem like a needless middleman, for the most part. We'll see in ten years. Because ten years ago it was self-driving cars around the corner, and now we have the models saying full chat and knowledge providers are around the corner. I'll be very surprised if the edge cases, where stuff actually matters are ever addressed beyond demo uses or marketing/advertising applications.
Looks like for Day 25, this would have marked the entire leaderboard (top 100) as suspicious. (No, they didn't use ChatGPT.)
Part of the meta-game while solving part one is to predict what kind of parameterization part two might depend on and make a flexible solution, and the best solvers are also the best at doing that.
I guess what I’m trying to figure out is whether anyone’s going to try and get a job using their Top 10 Advent of Code medal.
No one is going to get a job because of their AoC score, they’ll get a job because they have the knowledge and skill to complete the problems though.
The leaderboard fills up in anywhere from 10 to 15 minutes. I bet it would take people at least as long to come up with a suitable prompt (just copying the text on the site doesn’t work), refine the response, and implement the solution.