If you are using it to mix a snippet of code (from a sufficiently large code base) into a large code base of your own, then you are just remixing. That is not infringement. In music, there are entire genres based on remixing. You could even take it a step further and ask yourself: what is not a remix?
Point is songwriters absolutely get litigious over reproducing small portions of their IP.
complex.com/music/majority-sisqo-thong-song-publishing-owned-by-writer-livin-la-vida-loca
huffpost.com/entry/katy-perry-dark-horse-lawsuit-payment_n_5d43d825e4b0acb57fca3ff2
factmag.com/2016/06/25/sampling-hip-hop-copyright/
What's more surprising is to see copyleft advocates positioned so strongly in favour of giving copyright that kind of reach. I think that in a different context, some of the cases you refer to would be used by these same copyleft supporters as examples of why copyright needs to be more weakly enforced, not more strongly.
You have to filter out any non-copyrightable elements before you do the substantial similarity analysis. For code, that means removing non-expressive elements like arithmetic or boolean expressions, looping, recursion, conditionals, etc. APIs are not copyrightable under the recent Supreme Court holding in Google v. Oracle.
How much of your code is actually left after filtration?
But perhaps there could be a way to make something that automatically converts these "non-expressive elements" into copyrightable elements?
Output would probably be maddening to figure out though.
We could remove the non-copyrightable parts of text works too. Just take out all the basic building blocks of language like verbs, nouns, stop words...
The tests exist as a way of determining if someone did the action of copying which is the important thing at the core of it. And in this case the facts aren’t really in dispute, it’s whether what GH is doing counts as copying.
If you had a really really good memory and remembered almost exactly how your former company implemented something and when faced with a similar problem unknowingly produced similar code. It’s iffy as to whether this is copying. Because it really probably isn’t — you’re allowed to learn from copyrighted works — but the courts aren’t omniscient and when presented with the code they very well may rule that copying was more likely than not.
But we are omniscient in this case so we don’t really need the tests. Is what GH does more like copying or learning? This isn’t something that can be determined purely from the output of the tool.