It's definitely more complicated than just checking /var/mail but with the full set of features enabled you can get visibility:
- If the task fails to be placed, or stops with a non zero exit code CloudWatch event rule for the task event triggers allowing you to respond to that as you see fit: email someone? trigger a PagerDuty? - If the task logging was enabled you can go to the CloudWatch log stream for the task and see the actual stdout from the container and see why it had crashed.
I tried tweeting @AWSSupport, but they just referred me to the scheduled tasks docs that I had already read.
I suspect the issue is that I set up my tasks to run in Fargate mode, which I didn't realize at the time was brand new. Maybe it's not compatible with scheduled tasks yet.
For my “worker” ECS containers, I set them up to run their tasks using an app-level scheduling library. When a task finishes, the app phones home an event to Datadog. If Datadog doesn’t get notified within a certain amount of time, I get an email. So far it’s proven to be pretty reliable.
I’m sure similar functionality exists somewhere in AWS, but DD is so easy and we use it for a bunch of other stuff anyway.
After everything is up and running, all you need to do is submit a job (from a lambda function, CLI, whatever works best for you).
Then, in cloudwatch event rules, you can setup a new rule that only triggers on failed jobs, which can then have a SNS/Lambda target.
Or, alternatively, setup a cloudwatch alarm on the actual event rule invocations.