JobRunr: A library for background processing in Java
jobrunr.io
jobrunr.io
First of all, thanks to @mooreds to post JobRunr on HackerNews.
Second of all - I read some claims that being in the 'job scheduling' business is easy money. I would like to point out that's not really the case.
With JobRunr being open-source and more successful than I ever could imagine, this brings along a lot of stress. If you make a mistake (which I did in V6) the whole world starts to see it. I also try to keep the amount of open issues really small as these things linger in my head and also give me stress.
Anyway, this to say that I'm now able to provide my family with food but I'm still not break even (meaning if I just had freelanced as before, I would have more money in my bank account).
But, I can now work on something I love.
P.s.: it's indeed LGPL but this is also the case for hibernate. It means you should only open-source if you're touching part of the JobRunr code, not if you just use the lib.
See also https://github.com/jobrunr/jobrunr/discussions/769.
Enjoying the ride for the moment - it's wilder than I thought. I must confess that all the visits on the website via HN gave me quite the adrenaline rush :-).
At the end of the day, I don't want more to run more dedicated boxes for yet another jobs systems. I just want to hand off a container to the ether and say "please run this container until it stops, and do this once an hour or once a day." I don't want to get alerts in the middle of the night telling me that the Quartz scheduler has had some esoteric failure, and I don't want Jobs A, B, and C to get killed because Job D started doing something dumb.
Having a nice UI is cool, but I would rather not have more servers and relational databases and Java-cron libraries that can do dumbness in the middle of the night.
No, they're often not. Many struggle to ever make a profit.
I love the splash page too: simple and to the point. They aren't saving the world with AI, they're just making better cron jobs.
Seems to be good money in job running software. See Sidekiq in the ruby world: https://codecodeship.com/blog/2023-04-14-mike-perham
That said, tech like Temporal and other workflow managers do exist now... But maybe most people won't choose them because jobrunr is just an `mvn install` away.
Being able to reliably fire-and-forget or schedule background tasks from a web app can be really powerful.
I tried it in 2020 and was not very happy with it.
Serialization was (is?) deeply embedded in the API and my use case didn't need that and it was a large burden with no upside in my application. Then there were fluent builders which instead of collecting parameters just executed on them and it made it highly order dependent with possible invalid states without any indication of why.
I'd love a lightweight job scheduler with metadata and a dashboard, but with less magic in the java world.
Maybe I should give it a go again, maybe it has changed.
I think it is interesting that job scheduling, dependency graphs, dirty refresh logic, mutually exclusive execution are all relevant in the same space.
I recently working on some Java code to schedule mutually exclusively across threads. I wanted to schedule A if B is not running and then A if B is not running. Then alternately run the two. I think it's traditionally solved with a lock in distributed systems.
I think there is inspiration from GUI update logic too and potential for cache invalidation or cache refresh logic in background job processing systems.
How does JobRunr persist jobs? If I schedule a background job from an inflight request handler, does the job get persisted if there is a crash?
So in your case, the crash would likely leave it in an open/running state if it was already picked up, at which point timeout/retry rules would kick in after a restart.
If the job wasn’t running yet, just queued, then it would be business as usual upon restart.
Don't bother your legal team.
We're not another SaaS company and we don't have access to your data.
I love it. (even if that's not quite enough to completely ignore legal in my org)https://github.com/hibernate/hibernate-orm/blob/main/lgpl.tx...
https://www.tldrlegal.com/license/gnu-lesser-general-public-...
Almost every company doing real time/stream processing at the scale of TB/hour or greater is doing so by relying on Apache Kafka, written in Java and Scala.
I've personally been a member of a team that wrote several services (network collectors, observability processing) in Go/Rust/JVM (we preferred Kotlin) in parallel for performance comparisons and found the JVM services to show much better throughput.
Your perspective seems quite outdated. Possibly from before Go or Rust even existed?
[1] https://adamdrake.com/command-line-tools-can-be-235x-faster-...
2. Hadoop is used for batch processing and it can be kind of slow, that's true. However the whole point of Hadoop is that you can operate on actually big data, not the 2GB database in the blog post. If it fits on one HDD then you don't need Hadoop and it will just slow you down!
Distributed big data processing systems need big data to actually be useful. Small data that fits on a single machine can also be processed on a single machine, which will always be faster than using a cluster with orchestration, distribution and network overhead.
... Java?
You can natively compile Java these days.