Have you tried Jev and compared it against your alternative classifiers? Should take half an hour to do that, then you’ll have your answer (or someone who already did it can tell you here).
For fun I also tried using it at my local LLM router (all of my prompts and responses go through it for personal analytics) to decide which model to route tasks to based complexity etc. Again, underwhelming for my purposes, and so I'm not using it.
Not sure what real world use case it's best at, but I agree that it's simple enough to implement it yourself and see if it fits the type of work you are doing.