Talk about reinventing the wheel!
We ended up writing 114 line typescript job scheduler that uses Mongodb as a store. Mongodb provides the atomicity of scheduling (find and modify). Beyond that our UI is Robo 3T. The job collection has 4 simple states.
Any process can write a job document with the earliest time the document can run. It is also scaled with the rest of our application: more instances, more jobs that can run simultaneously.
But... I can see in shops that don't have this complexity yet, Jenkins might be just fine.
Edit to add: We decided against using another data source like a Queue which is something we would have used a long time ago, but we're already at an infrastructure complexity point where it would not be worth it - we already have Elasticsearch, MongoDB and Mysql so adding a AWS queue or rabbit would be out of the question at this point unless it provided a massive functionality we would need for something else.
It sounds like you have a very simple collection of jobs running that lack run-time complexity like remote hosts being unavailable or preventing a service from being overloaded with requests after it comes online because you have an ever-growing list of tasks waiting.
The actual complexities of DAG-based workflow management tools are considerably more complex than can be expressed in 114 lines of any language, and I hope any programmer I work with would spare me and my peers the misery of trying to roll and maintain our own when plenty of open source and actively maintained projects fit the bill that could benefit from our contributions.
I don't know why you would mention this because you don't have the requirements I have. If I had solved all scheduling needs for all people in 114 lines of code I would have written so.
Jenkins is leaky, unstable, unfriendly to the filesystem, buggy and poorly documented and tested.
If you’re going to do anything like this it’s actually worth paying for team city or something else.
In a gig last year, I needed to improve the build turnaround times on a Jenkins system. After learning everything I could about Jenkins, I realized the correct answer was to delete it and rewrite my own version that ran on the local system, which was way, way faster and much easier to debug and maintain.
Not having to commit/upload your code to a build server and then wait to get an executable/package back is an enormous time-saver just in that overhead alone, but even the build itself was faster, even though it was written entirely in bashscript.
Go figure.
"Maybe next year we'll have time to look into your plugin."
In any system I manage the cost of the plugin simply isn't an issue - if it's worth the money, it's easy to justify paying. But if it's free, and it becomes critical to the workflow, and in a year's time the author has abandoned it, then the "cost" is that we assume the responsibility for maintaining it forever, or we retool the workflow, or in some other way, it's very expensive. So paradoxically paid-for plugins are a much easier sell and requests for free stuff are shot down straight away.
Nothing you go on to detail after that opening sentence has anything to do with Jenkins itself but rather your company and it's organization.
Builds are just another job, usually triggered by a web hook after a git commit or periodic polling.
At my last job, we took cron jobs that were spread across 27 servers that we're really well monitored and moved them to Jenkins. The server SSH'd into the boxes (where necessary) ran the jobs, alerted us if there was a problem in Slack and we could use the tracked logs to find out exactly what happened. Really helpful.
Plus the cron scheduler has an alternative to picking a specific time to run by deferring to Jenkins based on other things that are running and the time it historically takes to run the job.
So instead of a bunch of jobs running at 3am you can set a job to run at [1-5]H and it will run sometime between 1-5am as best determined by Jenkins.
You can also trigger other jobs following the success (or failure) of another. For example, we used to have a daily digest that ran at the same time everyday and backed up our workers for a couple of hours. Instead, we created worker servers that only listened to that specific queue and scheduled a Jenkins job to start & update however many we needed, then after it was successful to trigger the digest jobs. Later in the evening, it was scheduled to scale down.
Just having all the jobs in one place with job specific log rotation rules and success tracking was extremely helpful.
It wasn't pretty...but it worked really, really well.
I think this is a case of everything looking like a nail to someone with a hammer. Just because you need a solution for scheduling and jenkins does scheduling doesn't make it the right tool for the job.
Hopefully everyone recommending it is at least running separate instances for production environments? Some (thankfully) former colleagues of mine had our build server running production jobs...
Everything is fully Dockerized, and Jenkins is set to just run the docker image with a specified entry point.
Different requirements, different solutions.