I started going down this rabbit hole a year ago (see the many good replies to this https://twitter.com/swyx/status/1241482183472295939?s=20) and most people feel it is "hard to reason about", which often seems code for unfamiliar.
What I and other people were lacking is a good framework to think about it. Unfortunately this has compromised my credibility to you as I left Amazon to go work on this very problem at https://temporal.io this year. I'll try to give some thoughts for how we tackle this but wanted to give that disclaimer upfront - not trying to sell you anything other than "i think this architecture could work use whatever you want"
1. DLQs - the AWS answer would be to wire up Lambda and SQS to build your own DLQ retry system (https://aws.amazon.com/blogs/compute/using-amazon-sqs-dead-l...). This is a bunch of extra provisioning and coding. So you may want to use the retries built into Step Functions (https://aws.amazon.com/blogs/developer/handling-errors-retri...). But instead of learning a bespoke States Language and debugging-by-redeploying-cloudformation (sooo slow lol), you may wish to work in a proper programming language SDK you can run and test locally instead (this is Temporal.io's approach)
2. Failures - any decent workflow engine will log and retry your failures for you, i wouldn't write my own logic for that these days
3. Microservice communication - what problems do you foresee? need more here. We simply call them Signals (send data in) and Queries (get data out) and it works well.
4. Breaking schema changes (versioning/migration) - yes this is really fragile unless you have a proper framework to bring this all together. We just build in versioning into our SDKs and give you a replay tools to verify you've handled still-running workflows (https://www.youtube.com/watch?v=kkP899WxgzY)
5. Keeping engineers happy - this one REALLY depends what youre talking about but being able to write tests for your asynchronous/distributed system is important for increasing confidence, as is being able to work in your preferred language (polyglot microservices), making every part of the system horizontally scalable so you don't have random bottlenecks, having everything logged and persisted so you are resistant to network/machine failures and can figure out exactly what went wrong when it goes wrong... I could go on.
Of course i'd love for more neutral users of workflow engines to chime in if I got anything wrong here. just trying to offer what I've learned so far working in this area.