Cool, so, looks like the product generates daily reports for advertisers about performance of ad spots over a trailing 1 day, lining up info about (a) what ad ran in what markets at what times plus (b) what increase in mobile/web traffic/conversion events happened from those markets coincidental with advertising, if I understand
https://www.decidata.tv/advertisers.html right. Looks like a complementary product (maybe running as a second process on the same nodes, or maybe running as part of the logic of the thing that parses the video) also produces quality reports about aired content (
https://www.decidata.tv/operators.html ). I don't really know video very well, probably some very clever stuff you can do here to parallelize the processing.
So hmmmm, a few things to think about for the tooling:
() How quickly can you notice a single-node service outage and react / restore service?
() How much downtime is ok? Is it ever ok to fail to process data for an entire day for one of the ~50 markets (can you just say "this report is incomplete" and reprocess it later)? How much wiggle room do you have (how long does it take 1 instance of the application to do its local processing job on 24 hours of data, and how long does the downline data pipeline have?)
() Are you responsible for the part that watches TV + saves it to disk, and ALSO the part that batch-processes the saved video and produces the ad metadata / quality metadata? Are those different services deployed separately?
() Yikes, is there, like, a person on call in each of 50 places who can unplug the thing that's watching TV and plug it back in, if a wire burns out or if the machine that's watching TV fails? That sounds like it could be an operational nightmare but maybe this is a solved problem / maybe you can buy watch-TV-and-save-it-to-disk as a service?
() How do you make sure that every node is running the same version of the software with the same version of all the configuration?
() How do you release new versions of the software? Do you (a) have a non-production environment which runs side-by-side at every node, processing the same data, and measure that the output is equally correct or better, the performance is the same or better, and there are no new errors or failures -- then promote the non-production code to production automatically after a while? Or do you (b) deploy a "canary" to one of the 50 nodes and watch how it performs for a while, then deploy to the rest if they all behave okay? Or do you (c) just ship the latest code to all 50 nodes on Friday nights when no one's looking, and check some health metrics over the weekend, and roll it back on Monday if a metric looks bad?
() How do you keep track of what software versions you released when? If someone asks "What changed on 'X date'" does your tooling let you tell them pretty quickly and pretty accurately?
() If you discover a bad bug (the parser is corrupting data; the new parsing job has a slow memory leak and fails catastrophically after 21 days of uptime) and you need to make a change in a hurry, do your runbooks let you make that kind of change quickly?
Doesn't seem like you really need k8s or ansible or chef or whatever for this, as much as you need to write down your operational requirements and decide what technology will meet them.