The original Nabla article is missing information on how they primed GPT-3 for each use-case, and how much effort they put into finding good ways of priming.
All fancy GPT-3 demos seem to rely on good priming.
The time scheduling problems are probably hard limit of GPT-3 capabilities. The "kill yourself" advice, on the other hand, might have been avoided by better priming.