Stackful coroutines are not new - it might surprise you to know that they are older than async/await, existing in full form probably introduced at least in 1967, with the first documented implementations starting around 1958. Stackful coroutines were implemented in a wide array of popular languages in the 1980s as pjmlp noted.
Stackless coroutines need better clarification, since the paper that defined this term [2] refers to a limited subset of stackless coroutines that we would nowadays just call "generators". To the best of my knowledge, when the paper came out generators were the only type of stackless coroutine in existence, and thus came the perception that stackless coroutines cannot be nested (they can) and that stackful coroutines are strictly superior to stackless coroutines (they aren't). You can see an example discussion here: https://news.ycombinator.com/item?id=16318535
The real difference is not nesting, but the fact that stackful coroutines manage a growable stack dynamically at runtime, while stackless coroutines pass the job on to the compiler to evaluate all possible nesting permutations, and create an optimally memory-efficient stack machine that represents the coroutine.
To summarize the pros and cons of each, I think you can say:
Stackful pros:
* [No colored functions](https://journal.stuffwithstuff.com/2015/02/01/what-color-is-...) - summed up it says: I can write asynchronous I/O code the same way I write multi-threaded code or non-blocking code. It's definitely more ergonomic, but also means your code can hold more surprises for you. Not everybody thinks 'Colored Functions' are an absolute advantage, otherwise we wouldn't have Effect Typing. * Less GC allocations: stack is [re-]allocated in bulk every time time a lightweight starts (go statement) or whenever it needs to grow. Stackless coroutines typically allocate a smaller object on every call to an async function.
Stackless pros: * Clear yield points: you always know the points where your coroutine could suspend and control would be moved to another coroutine: whenever you see an await statement[3]. This is why function colors are needed - to be explicit about the control flow. * Less memory waste and less overall allocation size: the "compile-time stacks" generated by stackless coroutines are perfectly efficient. There is no unnecessary memory allocated.
I don't think it is clear that one model is strictly superior to the other, but I do believe that if you're aiming for ease of correct use (rather than ease of use), the stackless model is clearly better. It is definitely somewhat harder to wrap your mind around, but it is also somewhat harder to get a deadlock or have a hard-to-detect data race in this model.
[1] Simula 67 (and probably Simula I too), and first implemented in Assembly as early as 1958 by Melvin Conway of Conway's Law fame.
[2] http://www.inf.puc-rio.br/~roberto/docs/MCC15-04.pdf by Roberto Ierusalimschy, the author of Lua.
[3] It's slightly different in an awaitless language like Kotlin - here you actually have to check whether the function you're calling is suspendable or not.