This part stands out:
"our team spent most of time monitoring jobs and and waiting for them to finish. Our solution for second place on the final leaderboard required 1 hour on 2500 CPUs"
Before I got to this part, I had assumed using AutoML would involve only reformatting the training/validation data, and then letting a single job run its course. Why does something that's 'automatic' need people to run multiple jobs?
Anyone know why they used CPUs instead of GPUs/TPUs? If they're distributing the computation over 100s of CPUs, then it's clear the computations can be done in parallel.