If I'm training codegen models, why wouldn't I just exclude code that contains these keywords? Shouldn't you have secret keywords, that people have to register to you, but you don't make public until after the fact, in order to avoid this?
For codegen, the results will always be only superficially useful. If AI could write code for us going forwards, it would imply there is a sufficient corpus of existing code from which to write remaining software. This is an astronomical miscalculation that fails to comprehend the vast complexity of program variations.
How sufficient is the existing body of code, compared to the code we might possibly choose to write? We can enumerate programs as tuples of sets of input,output pairs. So one program might produce 1 when you feed it 0, ie ((0,1)). Another might be represented as ((0,1),(123,456)) and so on. How many possible programs are there that transform trivial datatypes like single ASCII characters? It's the powerset 2**128. How many possible programs involve character pairs? 2**16384. These are numbers that make all the programs written to date look infinitesimal.
AI writing our code for us? AI a system that recycles our existing ridiculously tiny body of software to extrapolate what we might want to write, is not at all in the realm of possibility for what we are calling AI. GPT-4, as great as it is, is Google 2.0. That's it. The claims of 'AI writing my app' are just click bait.