It sounds like we're using very similar setups for writing. :)
I've primarily been using GPT-3 (and burning through millions of tokens) so I've been experimenting with GPT-J more lately and I've found it makes significantly more basic logic errors (e.g. mixing up pronouns, "forgetting" characters, using new characters), which makes me lean more towards constantly regenerating text versus revising it, and I feel like I'm backed into a box of having to babysit ~100-200 tokens at a time instead of generating significantly more at a time like I do with GPT-3.
I also built a quick tool that lets me adjust how much context I'm using in the prompt and generate 2-3 side-by-side completions to pick and choose from (just to speed up the flow of `click best suggestion` --> `keep generating from there`), but I haven't integrated GPT-J yet since it just feels... lower quality (it feels similar to ~Babbage, IMO).
But being comparatively* free, I'm still excited about GPT-J. Do you have any tips or processes you've found to make it spit out higher-quality text? 20k tokens at a time is quite a lot -- do you also have problems with winding paths / staying on a general "plot"?
Would love to hear any suggestions you have, because I'd sure love to move off GPT-3 to something comparable in quality!