- application startup and shutdown (at INFO level)
- occurrences that indicate a bug in the system logging (at ERROR level)
- occurrences that indicate a bug in another of your systems (at WARN level)
The last two require a lot of education and discipline to stop people from logging rare but valid circumstances, bad input originating outside your software, and "important" events. The answer to "but don't we want to know when..." arguments should always be, "If it's that important, emit it as a metric or store it in an appropriate datastore." Anything logged at WARN or ERROR should be something that can be addressed by fixing code.
If you stick to this discipline, monitoring logs is valuable and reasonably simple. You can alert on any WARN and ERROR level logging. There's nothing like looking at a month of logs and seeing only actionable information and a handful of application lifecycle events. (There's also nothing like looking at six months of logs from a stable system and only seeing lifecycle events.)
It is hard to stick to this discipline. You need a manager who believes in it and will back it up when engineers want to depart from it, and who will always prioritize fixing violations.
Why is it so hard to make engineers follow these rules? Because initially, it makes them feel incredibly anxious. It feels irresponsible to throw away so much "important" information instead of logging it. After working this way for a while, though, they realize that it actually forces a higher level of responsibility. Logging information in noisy logs might feel better than throwing it away, but it's not. It creates an illusion of having handled the information responsibly. By taking away the illusory option, you force people to make a real decision between throwing it away and saving it in an actionable way, by emitting it as a metric or storing it in a datastore to be processed.