> but I wonder why move away from the concept of metrics extracted from logs to discarding logs and storing just the metrics?
This is a _very_ good question.
Getting from log to meaningful dashboard full of metrics is very hard to do well. Its also very hard for a normal person(ie non programmer) to ask the right question of a log stream.
For example if you give me a log stream of a web server, and want me to raise an alert when an instance's health check takes longer than n milliseconds, it would require regex to get the host name, regex to get the health check, and more regex to get the response time. All of those steps are hard and require testing. Moreover its fragile, any kind of format change is really easy to slip in and cause your monitoring to break.
So your next question might be, why replace splunk with a metric based solution?
Thats the thing, you don't! logs are really really useful for what specifically went wrong. But they are not very good at telling you what is _currently_ going on over n number of services.
when implementing a service/program/things, I ask the developers to think about the information that is useful to log. Be that response time, number of peers connected, cache size, etc etc. Then instead of just dumping them to logs, you'd push them to a metric library and get that to look after it.
this means that you are consciously thinking about an number that best sums up the thing you care about. This then frees up logging to explain _why_ something has happened, rather than being rigidly defined because you'll break a dashboard.
Think of is as this:
dashboard: shows you if a service is happy, which is an aggregate of many smaller programmes.
Something goes wrong, response time of a microservice is above limits, you then track down when exactly the metric went wrong, and fire up splunk to look at the logs.