HNHacker News
TopNewBestAskShowJobs

firebacon

338 karma · joined April 29, 2018

submissionscomments
firebacon··on Elixir at PagerDuty
The types of values and legal conversions between them are well defined in C -- the assumption that everything is just a block of bytes and hence may be casted freely between types is explicitly incorrect (and is what may lead to bad/invalid C code).

However, C does allow you to write code that can not proven to be legal with respect to the formal system of the language and get it to compile, I think that is what you are referring to.

But just because you can write an invalid program does not mean the type system doesn't exist or is flawed. I can also write an invalid Haskell program.

The difference between these two would be that the haskell program can be proven to contain only operations that map to well-defined operations in the formal language model by an automated process, while the same can not be done for C in the general case.

However, just because you can not always prove conformance to a formal model using automated means, doesn't mean that no formal model exists (because it does!).

firebacon··on Elixir at PagerDuty
If we are being pedantic, what is going on here is not type coercion.

If it were type coercion, we would expect the add function to have either of these two behaviours:

  - coerce the input arguments to strings, i.e. implicitly accept two strings and return another string

  - coerce the input arguments to integers, i.e. implicitly accept two integers and return an integer
However, that is not what is happening here.

Of course, in Javascript it is implemented as a function that takes the theoretical "Any" type and decides what do do at runtime. But in a statically type language, the equivalent mechanism would be dynamic/multiple dispatch (or a static overload, possibly in combination with a generic/template method) -- not type coercion.

firebacon··on Elixir at PagerDuty
The result add(1: Int, "2": String) -> "12": String can be achieved in any language, regardless of how strict the type system is. All your example shows is that in elxir no such `add(int, string)` operation is provided as part of the standard library/environment, while in javascript, it is.

The "+" (operator) case is only special in languages that do not implement operator overloading -- it is completely orthogonal to the type system.

firebacon··on Elixir at PagerDuty
> Does anyone make weakly typed languages anymore?

I don't think anybody ever did. The definition of "strong" and "weak" seem to be entirely subjective; and of course everybody considers their favourite language to be "strongly typed". It's like a reverse no-true-scotsman.

firebacon··on Four years after its release, Kubernetes has come a long way
> Were there tools that helped keep all the services running together in a single machine?

Traditionally, services on unix system are started and supervised directly by the init process ("init system"). That has been SysV init for most of the time, but in the last decade we got two major new systems: upstart (dead now) and the infamous systemd. Using systemd can still be a very good alternative to using docker today.

> Was it a machine stack per service?

That still sounds like a good approach to me today.

If your service is not big/expensive enough to need at least one full machine, why bother cutting your app up into such small services in the first place?

For example, for web applications, a traditional setup would have been to have a couple of dedicated servers that run the database and a larger pool of webservers that serve the (stateless) application. In front of that you could have two linux boxes doing the network load balancing and failover or a commercial load balancer "box" product.

firebacon··on Four years after its release, Kubernetes has come a long way
That is a very interesting counterpoint! I guess I agree with you if the choice has to be between {Google, Amazon} or "only Amazon". But I'm advocating for the third choice, where we continue (or go back) to running our own infrastructure.
firebacon··on Four years after its release, Kubernetes has come a long way
Don't think the term "container" is really well-defined.

The "container" that docker and others implement is actually a collection of different kernel namespacing features. I assume the one you are referring to are cgroups. I think a better description would be that each process in a linux system is part of (many) cgroup hierarchies. And you can have more than one process in each of the groups.

I think what parent meant is that you can actually get all of these really nice isolation features for your service without using "Docker". It is trivial to enable them using linux command line utils, or use something like systemd which can also do it for you.

firebacon··on Four years after its release, Kubernetes has come a long way
You don't have to prove anything. Nobody is on trial here, I hope :)
firebacon··on Four years after its release, Kubernetes has come a long way
> Which part of kubernetes is "developed with the intention to eventually get you to use the hosted version"?

I believe it's all of it. Why else would they spend money on building and promoting it?

firebacon··on Four years after its release, Kubernetes has come a long way
Well I agree that the interests are aligned insofar as Google is trying do dethrone AWS which currently has a near-monopoly on cloud. But I'd still rather live in a world where the compute fabric of the internet is not centralized in the hand of a few corporations. I'd much rather get paid a salary from money that would have otherwise gone to a monthly *-Cloud payment. And I believe that in a lot of cases, the non-Cloud version of the product would actually be a better, cheaper, more reliable solution. But that will become a harder and harder sell the more successful the marketing from cloud providers is.
firebacon··on Four years after its release, Kubernetes has come a long way
Well yes, you can replace Google with like two or three other companies in my comment. At the end of the day I don't like the trend of outsourcing all the interesting bits of technology to a handful of mega corporations. I think it is not a good bargaining position for us developers, Google and the like can already get away with paying shitty salaries and it will only get worse the more control they have. For all the smaller companies, they are paying a cloud provider instead of hiring local talent to build solutions. It's loose/loose.
firebacon··on Four years after its release, Kubernetes has come a long way
I agree that running kubernetes ourselves does actually serve our own interests. But I still think it's much more of a slippery slope than alternative solutions (being that some of them are not developed with the intention to eventually get you to use the hosted version)
firebacon··on Four years after its release, Kubernetes has come a long way
Well I am sorry if you felt the wording of my comments was too "aggressive"; personal angle is weird though. Also I think you are trying pretty hard to misunderstand what I am trying to say. Obviously, using google's hosted version will make all your problems go "poof" at any scale, because they are Google's problems now. But that is precisely the point I am trying to make: It's not a feature of using kubernetes - it's a feature of paying Google!
firebacon··on Four years after its release, Kubernetes has come a long way
I think the kubernetes project is heavily driven by Google marketing, and that they are not doing this out of charity, but because they are trying to get you to use their cloud platform in the long run. They know getting somebody to build their stuff on open-source kubernetes is a win for them. After some time people will realize that running kubernetes yourself is actually harder and more fragile than just running your app without it and at that point the obvious move will be to use a hosted kubernetes service, like Google's cloud.

And I really think just handing over all our apps to google to run them for us is not in our (developers) interest in the long term. It would further solidify Google's "monopoly of the internet" position and also means that in the future - once we have succeeded in convincing our bosses to just rent every interesting bit of technology from google - that the only interesting jobs left will be... at Google.

So please go ahead and downvote me, but please also try to consider my point of view (that there is a ton of very aggressive marketing with a financial incentive going on here) next time you read and defend some kubernetes hype piece.

firebacon··on Four years after its release, Kubernetes has come a long way
I'm not sure I understand this. You have a product that is split into many different components, and when you deploy this product to a customer site, each component runs on different hosts, so you have a bunch of wiring up of service addresses to do for every deployment?

Could something like mDNS be a lightweight solution to that problem?

And also I am genuinely curious how kubernetes would solve that. When you install kubernetes on all of these machines, don't you have to manually do the configuration for that either? So isn't it just choosing to rather configure kubernetes instead of your own application for each deployment? If it is that much simpler to setup a kubernetes cluster than your app, maybe the solution is to put some effort into the configuration management part of your product?

firebacon··on Four years after its release, Kubernetes has come a long way
No it doesn't. You can not extrapolate from the fact that the hosted version "just works" that kubernetes at scale would also "just work" (kubernetes being the open source product that you run yourself here). Especially if the hosted version is offered by google and everybody knows that they employ truckloads of very good engineers keeping their stuff running.
firebacon··on Four years after its release, Kubernetes has come a long way
How does this relate to my point that you're transitively paying somebody else to do ops for you? Maybe google's pricing model rolls this into the normal VM price? Or maybe it's currently offered at a loss to gain traction? Or are you saying they are not paying the SREs anymore and they work for free now?
firebacon··on Four years after its release, Kubernetes has come a long way
The difference in simplicity is not in the interface that is presented to you as a user. The difference is that your shell script will have a couple hundred lines of code, while the docker and kubectl commands from above will pull in literally hundreds of thousands of lines of additional code (and therefore complexity) into your system.

I'm not saying that is a bad thing by itself, but there definitely is huge amount of added complexity behind the scenes.

firebacon··on Four years after its release, Kubernetes has come a long way
> But the parent's comment is missing the point

That was my point. I wanted to point out that while some people have only/first heard about failover in the context of kubernetes, it is not something that is specific to kubernetes or even the problem that kubernetes was build to solve.

Of course it is not designed to be a failover solution specifically and using it (exclusively) as such would be ill-advised; I was just trying to be diplomatic while pointing that out.

firebacon··on Four years after its release, Kubernetes has come a long way
You mean a failover solution? It's hard to give a list here because it is such a large space and it completely depends on your product/application. It is more like a category of use cases than one specific use case.

Some ideas for things you could do in a web/website/internet context assuming you have a single point of presence:

One type of "HA" is network-level failover; haproxy (L7), nginx (L7) and pacemaker (usually L3) seem to be very popular options, but I think there are dozens of other alternatives. In terms of network-level failover, things get more interesting once you are running in multiple locations, have more than one physical uplink to the internet or do the routing for your IP block yourself.

For application-level failover and HA, one option for client/server style applications is to move all state into a database which has some form of HA/failover built in (so pretty much every DB). I think this is very common for web applications and also for some intranet applications.

Assuming you have a more complex application running in a datacenter, there is also a lot of very interesting stuff you can do using a lock service like Zookeeper or etcd or really almost any database. Of course, you can also make your app highly available without using an external service; there is a mountain of failover/HA research and algorithms that you can put directly into your app (2pc, paxos, raft, etc). Of course all of these require some cooperation/knowledge from the application. For some apps it might be very hard to make them "HA" without relying on an external system, but for some apps it will be trivial.

Note that when you move away from a web/datacenter context, to something like telecommunication or industrial automation, so something that doesn't run as a process in a datacenter but is implemented {as, on} hardware in the field, failover and high availability will have an entirely different meaning and will be done in a totally different way.

firebacon··on Four years after its release, Kubernetes has come a long way
> makes economic sense [...] for our current team.

But the reason for that is not that it makes "hard problems go poof at scale". The reason is that you're using a hosted service where somebody else (in this case, Google themselves) takes care of the problems for you for a fee and - at small scale - you only have to pay them a fraction of a single operation's engineers salary for it.

So of course it makes economic sense for you to use a hosted service where the sharing economy kicks in, but your recommendation to use kubernetes because it solves hard technical problems at medium scale does not follow from that.

firebacon··on Four years after its release, Kubernetes has come a long way
Nobody wants to wake up at 1AM because their singly-homed service just went down. Kubernetes might be a fine tool to achieve that, but I want to point out that there are much simpler failover solutions available. Failover is something you should definitely have, but also something you definitely don't need kubernetes for.
firebacon··on Four years after its release, Kubernetes has come a long way
When you notice that you spend such significant money (i.e. multiple salaries) on hardware capacity planning that employing a team to operate kubernetes might be cheaper instead. Probably not going to happen unless you run on hundreds to thousands of machines.
firebacon··on Four years after its release, Kubernetes has come a long way
> because a lot of hard problems at medium scale and above just go poof with K8s.

No, they don't? I don't know why anybody would just assume that something as complex as kubernetes would just run flawlessly once you actually try to run it on thousands of servers. Must be something to do with google PR because people definitely don't seem to assume the same for e.g. hadoop or openstack. Make a guess at how many people large companies have to employ to actually keep their smart cluster scheduler running?

firebacon··on Four years after its release, Kubernetes has come a long way
I think the recommendation is not to use another orchestration service, but to keep the infrastructure simple and not use any orchestration or service discovery at all.

Adding service discovery and container orchestration will probably not make your product better. Instead it will add more moving parts that can fail to your system and make operations more complex. So IMO a "containerized microservice architecture" is not a "feature" that you should add to your stack just because. It is a feature you should add to your stack once the benefits outweigh the costs, which IMO only happens at huge scale.

Most people know that "Google {does, invented} containers". What not so many developers seem to realize is that a Google Borg "container" is a fundamentally different thing from a Docker container. Borg containers at Google are not really a dependency management mechanism; they solve the problem of scheduling heterogenous workloads developed by tens of thousands of engineers on the same shared compute infrastructure. This however, is a problem that most companies simply do not have, as they are just not running at the required scale. At reasonable scale, buying a bunch of servers will always be cheaper than employing a Kubernetes team.

And if you do decide to add a clever cluster scheduler to your system it will probably not improve reliability, but will actually do the opposite. Even something like borg is not a panacea; you will occasionally have weird interactions from two incompatible workloads that happen to get scheduled on the same hardware, i.e. operational problems you would not have had without the "clever" scheduler. So again, unless you need it, you shouldn't use it.

I think the problem that Docker does solve for small startups is that it gives you a a repeatable and portable environment. It makes it easy to create an image of your software that you are sure will run on the target server without having talk to the sysadmins or ops departments. But you don't need kubernetes to get that. And while I can appreciate this benefit of using docker, I still think shipping a whole linux environment for every application is not the right long-term way to do it. It does not really "solve" the problem of reliable deployments on linux; it is just a (pretty nice) workaround.

firebacon··on Showdown: MySQL 8 vs. PostgreSQL 10
I agree, if we are allowed to make the UUID arbitrarily large (and not be limited to the 128 bits from the "official" UUID algorithm), it should always be possible to set B to a large enough value so that our scheme assigns a unique and finite-sized UUID to every possible state of the system, since all the system's parameters (number of servers, time, etc) should be discrete and finite. I.e. basically make B large enough that 2^B becomes greater than the number of discrete timesteps in our experiment. That's an interesting observation.
firebacon··on Centrifuge: a reliable system for delivering billions of events per day
Ah ok; I have made a habit of typing "<name> {software,github}" into google as well as the WIPO trademark search when picking a name for a new project.

I started doing that after having to go through two iterations of renaming a project after finding out the old name was somehow encumbered... of course after we had already used it in front of other people; it was very embarrassing.

firebacon··on Centrifuge: a reliable system for delivering billions of events per day
Well I for one think that it's not really "cool" to just copy the name for more or less the same thing.

BTW, it also appears there is an active US trademark (#78010053) for "centrifuge" in the computer database context. Isn't that an issue, too?

firebacon··on Showdown: MySQL 8 vs. PostgreSQL 10
That only seems possible if both processes were running on the same physical machine (and hence share the same input MAC address).

It's a nasty corner case, but I'm sure the UUID designers considered it a valid tradeoff, since I believe the algorithm was designed for a distributed scenario. In the single-machine case an atomic counter would be a much easier solution with very reasonable efficiency anyway. Still, it might have been clever to also include the local process id in the UUID, I wonder why they didn't to that.

At any rate the problem is easily worked around by running the two processes on different machines, i.e. ensuring you have at most one UUID generating process per host (with respect to the database table in question).

firebacon··on Showdown: MySQL 8 vs. PostgreSQL 10
The difference here seems to be the exact definition of "unique".

A UUID has a finite length, so if we generate N new UUIDs in a finite time-interval it seems clear that - in theory - we must have collisions for large values of N (or even infinite N), regardless of how clever we are in seeding our random number generator. But I don't think this can be fixed without using a variable length value or refusing to mint new UUIDs once the available bits are used up.

Of course in practice that should not really happen; at least if we only run on hardware from vendors where we can assume that every MAC address will be unique, the only way to actually get a colission - in practice - would be by generating more than 2^B UUIDs on the same machine within a fairly short timeframe and fairly large B.

EDIT: In v1 of the algorithm it seems that B=14 and the "short timeframe" is the smallest resolution increment your system clock supports. That is if we assume the "uniquifying" clock sequence is produced by incrementing a counter. So we can say that for practical purposes, a collision is impossible on a modern system unless we have invalid MAC assignments to our hardware, unreasonable transaction rate, or an incorrect implementation of the UUID algorithm.

Am I missing something?

← PreviousPage 3 of 4Next →