HNHacker News
TopNewBestAskShowJobs

mmanciop

59 karma · joined September 4, 2021

submissionscomments
mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
Indeed if there were “official” container images out there, I might have instead run the server on Google Cloud Run or AWS AppRunner, without having to take care of the Linux underneath. Or an Amazon ECS task. I don’t have a Kubernetes cluster, but I will at some point make a version of this blog to run it on K8s :-)
mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
> Telemetry does not address this, though. Shoving it into a container and assigning it a simple “restart if down” rule does.

A Systemd unit as shown in [1] does it too without using containers and with fewer moving parts of using containers. I use containers every day at $work. I have been using containers since before Docker was a thing. In this case, it's entirely overkill: Systemd units use the important things like cgroups already.

For the upgrade: depends. You do need a container image regardless, and I have not seen official ones. Upgrading servers in Minecraft requires upgrading clients to match, and my kids prefer to play, more than upgrade. (Unless a biome is released. Then it must be immediately available to them.) But then again, I just need to download the binary with a cURL call. And if the configurations change, Docker won't help me there one bit anyhow.

[1] https://github.com/dash0hq/minecraft-server/blob/main/drople...

mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
To be fair, the setup of the article works with most modern observability solutions, in same cases just by replacing the endpoint and authentication token. Turning telemetry processing into a sort of utility is one of the great things that OpenTelemetry did. Now, among vendors, we compete on delivering insights on the telemetry, as opposed to just collecting it. If you are interested, I wrote about it a while back [1].

About excessive telemetry, that depends on what you want to achieve. Using facilities in the OpenTelemetry Collector like [2], you can easily drop all telemetry you have no use for. At the cost of tooting my own horn, we actually provide super easy ways of doing the same dropping at no charge whatsoever to the end user in Dash0 [3].

[1] https://thenewstack.io/is-otel-the-last-observability-agent-...

[2] https://github.com/open-telemetry/opentelemetry-collector-co...

[3] https://www.dash0.com/changelog/spam-filters

mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
Indeed.

My personal definition of nanosecond is the time passing between the Minecraft server having a hiccup, and the first scream piercing the air.

The printer not printing is DEFCON 5 material.

mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
I actually had to splurge got 2 VCPUs on Digital Ocean to avoid "skipping ticks" and it does sound pretty nuts to me. We play max 3 players. I would expect the server with such a load to be able to run on a slightly tuned up toaster.
mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
> As you said "There is so much fun that could be had in setting up an actually resilient system instead.", maybe the author has more fun setting up alerts and metrics instead of a resilient system like you do?

Adding the backup for the world files, already having Systemd bringing back a crashing server, makes the setup rather resilient. Sure, there's infinite more things that can go wrong, but with swiftly decreasing likelihood.

> The truth is that in most real-world scenarios getting alerts, metrics is much more important than building a fully resilient system (Expensive, maybe overengieering for early stage etc.).

This, very much this.

> However, the author is not really procrastinating—he gets paid for this. As the first sentence in the blog post says "One of the secret pleasures of life is to be paid for things you would do for free.", which I can very much understand as I often work or play with things I could use at work in my free time.

Yes :-)

mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
> How are metrics helpful? There is so much fun that could be had in setting up an actually resilient system instead.

Metrics are the means to an end of alerting. And with alerting, I mean getting pinged on my phone when something important breaks. Like, you know, the server going down.

> Why worry over metrics and alerts when you could orchestrate an infrastructure that grants you the superpower of being able to spin up a server with a copy of the world on a whim instead (or even a system that auto-starts one whenever there is demand)?

As somebody who has run cloud and enterprise software for almost two decades now, I can be that needs monitoring too. The more moving parts there are, the more things go wrong. The more things go wrong, and the more you care they get fixed, the more monitoring you need :-)

mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
I swear I had a lot of fun setting doing the setup.

I am also a massive observability nerd, so YMMV :-)

mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
One of the stretch goals for me writing this article was indeed to show between the lines how Prometheus Exporters, the OpenTelemetry Collector and Systemd can all work together. That is a very reusable skill on monitoring workflows running outside containers on Linux VMs or hosts.
mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
For this, I have the impression that https://github.com/dirien/minectl might be very close to what you are thinking. I did not try it, but took the Minecraft Exporter from it and used in the setup.
mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
Because I enjoy observability and monitoring a LOT, and because my kids nag me to hell and back when our home IT infrastructure is having a bad day.
mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
I am indeed an employee of Dash0. The setup for telemetry collector will work with anything that accepts OTLP, and with minor adjustments, the data can be sent elsewhere too in other formats, as the OpenTelemetry Collector is very flexible in that regard.

Alerting is specific to Dash0. I know of no other monitoring solution that lets you run real PromQL on logs. But there will be similar ways of accomplishing the same alerting logic.

mmanciop··on Monitoring my Minecraft server with OpenTelemetry and Prometheus
Thanks for setting me straight :-) I updated the article to reflect that.
mmanciop··on How to Inspect React Server Component Activity with Next.js and OpenTelemetry
For local development, yes. But I see that almost never done in production settings, especially with containerized workloads. Is there a neat way to do it in Kubernetes?
mmanciop··on Memory Allocation in Zig. Working with Netlink
I loved this article. I am working on an LD_PRELOAD object for some observability usecases, and while there was a learning curve, Zig an its lack of dependence from LibC is a breath of fresh air!
mmanciop··on I got OpenTelemetry to work. But why was it so complicated?
Adopting OpenTelemetry does not have to be hard for common use-cases. On Kubernetes, the Dash0 operator (https://artifacthub.io/packages/search?repo=dash0-operator) automatically instruments Node.js and Java workloads (and soon other runtimes) with just a custom resource created in a namespace. It works with all OpenTelemetry backends I know of.

Disclaimer: I am one of the authors of the Dash0 operator and work on Dash0 (https://www.dash0.com/), an OpenTelemetry-native observability platform.

Automatic instrumentation on Kubernetes is also provided by the community OpenTelemetry (https://github.com/open-telemetry/opentelemetry-operator).

I am certainly biased here because OpenTelemetry and Prometheus have been at the core of my professional life for the past half decade, but I think that the biggest challenge, is that there are many different ways to get you to a good setup, and people get lost in the discovery of the available options.

mmanciop··on Kubernetes horizontal pod autoscaling powered by an OpenTelemetry-native tool
Yes, it probably would be able to do the same logic as in the blogpost with the Prometheus scaler: https://keda.sh/docs/2.16/scalers/prometheus/

I might give it a go in a follow-up :-)

mmanciop··on Kubernetes horizontal pod autoscaling powered by an OpenTelemetry-native tool
True, but that approach is also limited in what it can autoscale on, namely Utilization metrics. They are not always good data to make scaling decisions on (covered in the article).
mmanciop··on Kubernetes horizontal pod autoscaling powered by an OpenTelemetry-native tool
I looked into how to wire the Horizontal Pod Autoscaler for Kubernetes to Dash0, the OpenTelemetry-native observability tool I am working on, and it turned out to be surprising simple.
mmanciop··on Show HN: Dash0 – Dev-Friendly OpenTelemetry Observability with Open Standards
We explicitly test 1.28+. But I have in front of me a 1.27.13 and the operator works on it without issues. And probably goes way further back than that.

What is the oldest version of Kubernetes you have seen in the wild recently?

Disclaimer: I work at Dash0.

mmanciop··on OpenTelemetry and vendor neutrality: how to build an observability strategy
I can 100% confirm that OpenTelemetry is a fantastic project to get rid of most observability lock-in.

For context: I am the Head of Product of Dash0, a recently-launched Observability product 100% based on OpenTelemetry. (And Dash0 is not even the first observability based on OpenTelemetry I work on.)

OTLP as a wire protocol goes a long way in ensuring that your telemetry can be ingested by a variety of vendors, and software like the OpenTelemetry Collector enables you to forward the same data to multiple backends at the same time.

Semantic conventions, when implemented correctly by the instrumentations, put the burden of "making telemetry look right" on the side of the vendor, and that is a fantastic development for the practice of observability.

However, of course, there is more to vendor lock-in than "can it ingest the same data". The two other biggest sources of lock in are:

1) Query languages: Vendors that use proprietary query languages lock your alerting rules and dashboards (and institutional knowledge!) behind them. There is no "official" OpenTelemetry query language, but at Dash0 we found that PromQL suffices to do all types of alerting and dashboards. (Yes, even for logs and traces!)

2) Integrations with your company processes, e.g., reporting or on-call.

mmanciop··on Ask HN: For those trying to hire how hard is it to find people currently?
And there I thought that being good at scamming was a key skill in crypto.

(Nothing personal, but it had to be said.)

mmanciop··on What OpenAI really wants
The just want tons of moneys
mmanciop··on A response to the git.centos.org changes
Don’t forget they also urgently needed to paint themselves as victims of the ecosystem they shaped and nurtured for decades.
mmanciop··on A response to the git.centos.org changes
While the post is mostly correct and one can agree with most of it, it’s very hard to believe that, unless more people pay for RHEL, it would cease to exist. What’s more believable, is that this is a matter of wanting more money for sales targets. Very legitimate and all, but don’t cry over the long hours and don’t try to sound like misunderstood heroes. RHEL is valuable, it contributes to the ecosystem well, and execs want it to churn out more cash.
mmanciop··on Microsoft Freezes Salaries for 2023
Translation: we fired a bunch of people, so did everybody. Talent’s no longer scare, and there is a large supply of jobless people that would get your position on a whim. Get fucked.
mmanciop··on Ask HN: Is the job market brutal? or is it just me?
> Lockheed is dirty while Palantir or Facebook aren’t

Said nobody with any common sense, ever. Palanthir or Meta are in my eyes as bad on CV as… as… wait is there a way I don’t go Godwin’s in this thread?

mmanciop··on Amazon rescinding the offer while I am on the notice period
Remorse? Are you kidding? Some likely get a long-forgotten tingle out of it ;-)
mmanciop··on Rackspace Exchange hosting down ~18 hours; “solution” is to start over on O365
With the mind-numbing level of incompetence that they are displaying here, would you actually consider availing yourself of their services? As the saying goes, fool me once, shame on you, fool me twice,…
mmanciop··on Ask HN: Why can’t cloud providers show the instantaneous consumption / cost?
And by “marginal utility of that information” I mean “marginal utility of that information being available in near real-time (anything below an hour in this context).”
Page 1 of 2Next →