The training dataset is not that huge (see https://github.com/metarank/ranklens/ for details, it's open-source), so we do a full retraining directly on the node right after the deployment, and it takes around 1 minute to finish. We also run the same process inside the CI: https://github.com/metarank/metarank/blob/master/run_e2e.sh
There is an option to run this thing in a distributed mode:
* training is done using a separate batch job running on Apache Flink (and on k8s using flink's integration)
* feature updates are done in a separate streaming Flink job, writing everything in Redis
* The API fetches latest feature values from Redis and runs the ML model.
The dev-mode I've mentioned earlier is when all these three things are bundled together in a single process to make it easier to play with the tool. But we didn't spent much time testing distributed setup, as this thing is still a hobby side-project and we're limited in time spent developing it.