When GTP-3 writes a scientific paper it's not trying to test a premise or critically evaluate some data, it's trying to generate text that looks like that sort of thing.
It doesn't matter how many excellent quality papers you fed it, or how stringently you excluded low quality input data, it would still only be trying to produce facsimiles. It wouldn't be actually trying to do the things a real conscientious scientist is trying to do when they write a paper. Arguably at best it might be trying to do what a deceitful scientist trying to get credit for a paper with spurious results with no actual scientific merit might be doing when writing a paper that looks plausible, but thats actually a completely different activity.
The code generation examples are really interesting. Here it's being used to generate real working code that has actual value, and it appears to work pretty well. It's code often has bugs, but who's doesn't? It only works for fairly short precisely definable coding tasks though. I don't think scaling it up to more complex coding problems is going to work. Again it doesn't understand the meaning of anything. It's trying to produce code that looks like working code, not actually solve the programming task you're giving it. It doesn't even know what a program is or what a programming task is. It doesn't know what input and output are. For example you can't ask it to modify existing code to change it's behaviour. It doesn't know code has behaviour. It has no idea what that even means and has no way to find out or any route to gaining that capability, because that's not a text transformation task and all it does is transform text.