You suggest up/down sampling rare cases. Can you please elaborate on the standard approaches for this kind of problem? For both linear and nonlinear classifiers. Thank you.
While I'm not an expert of the theory behind sampling, if you do find the need to tweak sampling to align the default loss function and the business metric, I would say doing grid search first, and validate the result with the business insight, e.g. if you find getting the rare cases right is much more important that getting the common cases right, does that align with the business insight?)