Innovating Cron: Announcing Norc
blog.perpetually.com
blog.perpetually.com
It was originally written in the context of Django but it is a general purpose Python library now.
How does it handle logs, resources and changing trees?
Gonna have to look into it more!
I don't think you can change trees after they are sent (you mean subtasks, right?).
You can see logging example at "defining and executing tasks" section here: http://ask.github.com/celery/introduction.html#usage
Sorry, I am only a light user at this point in time.
We'll be adding a RabbitMQ plugin to Norc, just like SQS, when we can.
"From our testing, we expect easily-achievable throughputs of 4000 persistent, non-transacted one-kilobyte messages per second (Intel Pentium D, 2.8GHz, dual core, gigabit ethernet) from a single RabbitMQ broker node writing to a single spindle."
Or what do you mean by serious?
Also, hmm, I didn't think rabbit was so significantly behind other AMQP implementations like zeromq: "4,100,000 messages a second" - http://www.zeromq.org/
Resources: There is AMQP QoS which makes sure it only receives as many tasks as it can handle.
Task hard and soft time-limits is coming in 1.0 (patch ready). If the soft timeout is exceeded an exception is raised which the task can catch to do any clean up before the hard time limit is exceeded and the task is forcefully killed.
Rate limit (per task type or global) using the token bucket algorithm (which allows for bursts of data). For 1.0 (patch ready and tested)
Otherwise you have to OS process resource limits (cpu/memory etc).
Monitoring is coming in 1.0 as well, someone is working on a monitoring system with a web-frontend where you can see the current state of the system (support for deleting already published tasks might be added, but then on an opt-in basis)
The current scheduling system is flawed (it uses the database, which is a dead end in my opinion), a new solution is almost ready which uses a separate centralized service that works like a clock sending out messages at schedule time: http://wiki.github.com/ask/celery/rewriting-the-periodic-tas...
Now for changing trees, I'm not sure what you mean here, please correct me if I misunderstood. Messages can not be changed once they have been published, so the task itself is responsible for changing the execution order. You can chain tasks, so say TaskA launches another task. You can retry tasks if they fail.
Oh and there's the message routing features made available by AMQP, which means you can have different servers/instances handle different tasks.
Celery has a lot of features, and even more is under development, so I don't think I can list them all here. I can't see anything hindering a Celery implementation of Norc, but as I read it you started working on this before celery started. Bad luck when we could have shared a lot of work :(
>You can chain tasks, so say TaskA launches another task. You can retry tasks if they fail.
How does this chaining work? Does each Task define its children? I've found that its cleaner to separate tasks from their place in the tree. For example, a script that downloads a CSV file each hour shouldn't care how that file is used.
The only annoyance is Apple's worst-of-both-worlds XML plist format, which replaced something fairly close to JSON with an awful pair-wise angle-bracket shit-pile.
If the goals are improved reliability, management, fault-tolerance, etc., you'd be better off using mature software that was designed to solve this problem, and already has good community support, like Condor.
It's like wrapper scripts. They work well until they don't: managing the logs, getting efficient execution, understanding what happens when all become quite unwieldy, whereas a Norc-like approach proves more self-documenting and easy to manage.
I believe resource management is limited to overall system load. Batch isn't designed for things like managing available licenses.
I'm not a huge expert at batch, so if I've missed something let me know.
Condor has DAGMan (http://www.cs.wisc.edu/condor/dagman/) for managing DAGs of jobs, and I suspect that it can scale higher than just a handful of tasks.
I believe resource management is limited to overall system load. Batch isn't designed for things like managing available licenses.
Condor acts as match maker for jobs and cluster nodes using "classified advertisements" (http://www.cs.wisc.edu/condor/classad/). Classads allow you to describe the jobs and the cluster nodes, and to express arbitrary requirements and preferences for both. This allows jobs to say "I need a node with X", but it also allows nodes to express requirements (e.g. "I won't any run jobs coming from the psych department"). I don't see why you couldn't use this system for expressing requirements regarding licenses.
However, Condor appears to be very complex. The PDF version of the Condor 7.3 manual has 991 pages in it. OTOH, Red Hat MRG Grid is based on Condor (http://www.redhat.com/mrg/grid/), so commercial support should be available.
People do just that: http://www.cs.wisc.edu/condor/techpaper/licenses.html
I would not say it is "complex" but ultra flexible. You will not have things working in just one day or anything, but I've set it up without a significant hassle. People have had a lot of ongoing success with Condor in large (and very large) computing environments.
Without that, I'd say it would be a slightly masochistic exercise to take a big system like condor and try to make it a cron replacement; while it clearly has that capability, if that's not one of its use cases, it's going to be a real bear to force it into that configuration. I've been a fringe user of condor before, and it can result in some unfriendly "rejected your job for unknown reasons" situations.
Right, why would a message queue support scheduling?
We use Norc in conjunction with SQS, and find the former is good for scheduling, management, logging, while SQS is better for lots and lots of repetitive tasks that we want processed ASAP.
Perpetually.com crawls, archives and versions any web site on demand, with any repeating schedule. One of the challenges is managing bursts of activity, so what we do is use Norc to manage the timing of archive requests, then SQS to manage farming out the actual archiving to available hosts. Since we hook SQS into Norc we can easily monitor/audit any system delay or outage.
We also use it to handle system backups, sending alerts and other system administration tasks.
While cron is great, it’s not geared toward solving this problem: Tasks are tied to a single computer, and they’re managed independently for each host and user from the command line.
That said, norc appears interesting, and I'll certainly be keeping an eye on it (I doubt I'll switch until it's a bit more established - and available in my favourite distros :)
In this case, yes, you can retry a task in the middle of a tree of tasks, and retry all the children tasks as well. This will leave some tasks untouched (the parents and peers) and retry the children in the same order as the first time.
If a task in the middle of the tree fails children tasks will not run until it's been skipped or successfully run.
Retrying children tasks would be a great job for a GUI for Norc because then you could choose which children tasks to retry interactively from a web app or somesuch...
I was little confused because it seemed like norc would only start jobs at the top level.
Jobs start with any tasks that return True when asked due_to_run(), which can be schedule, parents' status, etc.
Glad I could clear it up!
I know this is a bad place to ask for it, but we actually have just started a story at my company to look for a better scheduling tool--mostly, just a distributed cron. Do you have any preference for Q&A? Should we use HN?
If not, feel free to send me an email so I could bounce some questions off you (amcfague at wgen dot net)
EDIT: Considering these questions are not answered in the current README, I'd also like to be able to propose additions to the documentation as a result! :)