Consul.io – Service discovery and configuration made easy
consul.io
consul.io
It would be wonderful if all (most?) open-source projects have a such beautiful documentation explaining their architecture. Thank you Armon and team for sharing this with us!
http://www.serfdom.io/docs/internals/security.html
(If you're confused about how this relates to Consul, see this: http://www.consul.io/docs/internals/security.html)
Worse, their justification for using this instead of standard schemes like (D)TLS seems to be that they don't need the sorts of features transport encryption normally needs because they lean on the protocol's state machine to provide some of the features transport encryption would normally provide for free, like replay attack avoidance.
Not only does homebrewing encryption almost always carry with it design and implementation mistakes, but someone auditing the protocol not only needs to understand the cryptographic design, but how it interacts with the rest of the protocol's state machine.
My advice would be to scrap their homebrew encryption and use (D)TLS
We did develop the algorithm in the open (see this gist: https://gist.github.com/armon/7159161), and made a call on the community to provide feedback to help improve the design of the system.
The crypto system used is actually pretty vanilla based on AES-GCM. We certainly didn't attempt to invent our crypto primitives, and instead stuck to the best practices around the most modern systems.
Unless you mean you're doing multicast, you're still doing 1-to-1 connections.
> We did develop the algorithm in the open and made a call on the community to provide feedback to help improve the design of the system.
I wouldn't trust myself, let alone many unknown people who are just volunteering their time and have no known credentials as a cryptographer.
> instead stuck to the best practices around the most modern systems.
Except for the whole "don't role your own crypto system" one.
There are just so many things that can go wrong and when crypto fails it can do so very quietly. Additionally, if you ever want anyone to interoperate with you it'll just be a PITA for them, and depending on who it is, a PITA for you.
The overhead of a proper protocol like (D)TLS isn't that much.
> "On our production frontend machines, SSL/TLS accounts for less than 1% of the CPU load, less than 10 KB of memory per connection and less than 2% of network overhead. Many people believe that SSL/TLS takes a lot of CPU time and we hope the preceding numbers will help to dispel that." - Adam Langley, Google
I guess it depends on your definition of roll your own. We didn't invent AES-GCM or implement it. We are using the implementation shipped with the Golang stdlib.
This makes your protocol difficult to audit: someone concerned about potential attacks can't just look at your protocol in isolation, but has to factor the underlying protocol state machine into the security of your transport encryption protocol.
OK, so I'm not understanding why isn't a 1-to-1 message, nor why DTLS isn't an option here.
> I guess it depends on your definition of roll your own. We didn't invent AES-GCM or implement it. We are using the implementation shipped with the Golang stdlib.
You need to read better. I didn't say "crypto primitives", I've said "crypto _system_". That includes everything, including primitives, key management, authentication, replay (which means your application protocol is now part of the crypto system, not a good sign), field concatenation for singing/hashing, &c.
The most important thing is that when your improvised system fails, you will more-than-likely never know and it'll never cause any errors.
This isn't a particularly sophisticated cryptosystem --- AES-GCM with keys broadcasted between the trusted participants --- but it's simple and, while GCM is nobody's favorite AEAD, it is at least a formal AEAD mode, leaving not a whole lot of room to screw the basics up.
B) This works really well in conjunction with registrator https://github.com/progrium/registrator which listens to Docker events and will automatically register your services with Consul.
I occasionally run into an issue where a node will be briefly marked as failing. This seems to clear up pretty quickly but I can't really figure out what is happening when this occurs. The nodes are all in the same DC and on their own private network. I also occasionally lose all running Docker containers on a particular node. I don't think that is related to Consul, just throwing that out there :)
E.g. you got webservices talking over HTTP. One goes bad, you want existing clients not to talk to them.
Smart Stack: you are using HTTP proxy on each machine. You're code doesn't has to be aware of who is talking to.
Consul: Not sure, but it sounds like server code has to be able to handle events to change that.
With a tool like consul-template, you can have the same behavior that Synapse gives you. https://hashicorp.com/blog/introducing-consul-template.html
In the general case, clients of Consul don't need special handling code for events, it's handled by Consul for you.
Consul: your calling webservice uses a DNS name served by consul. Consul responds with a pool of working servers. As soon as one fails, consul doesn't include its address anymore, and hence the calling services stop talking to it.
However:
[1] https://www.consul.io/intro/vs/zookeeper.html
This did not unequivocally convince me in the benefits over ZooKeeper. On the contrary, it makes it seems that ZooKeeper tries to do less, and I strongly prefer simpler tools.
[2] http://aphyr.com/posts/316-call-me-maybe-etcd-and-consul
Have the issues discussed in this article been fully addressed yet?
Overall, can you outline a use case where Consul is definitely better than ZooKeeper?
With respect to ZooKeeper, there are different approaches. ZK provides a low-level primitive on which you can build. Consul provides similar primitives, but it ships with many features out of the box that don't require any development effort. It's a "batteries included" approach.
Specific examples:
- Real-time configuration with Consul + consul-template
- DNS based service discovery
- Scalable Nagios replacement
- Dynamic HAProxy / Varnish configuration
- Application configuration with Consul + envconsul
- Triggering config management tools with the event system
That is just a handful of uses for Consul that don't require writing any code. Doing similar things with ZK is possible, just requires a lot more work.
We don't have permission at this time to share some of our big users (we're working on approvals!), but I can say that that some very big companies you've certainly heard of and probably use right now have deployed Consul across every server. And Consuls deployment in just companies were working with is in the hundreds of thousands of machines.
At this point were very confident in its stability and it's well proven at very large scale.
If you're a big user and need references to other large users, email me and I'd be happy to set that up.
Without divilging too many details, what would be the ratio of agents running '-server' to overall agents?
All machines using consul will (generally) run the agent locally.
In general you are going to set up an N+M topology of n servers and n+m agents -- but for a network of 100k nodes does consul scale to 100k agents? Plus what goes unstated is how those 100k nodes are laid out in DCs and racks.
Anyway, this isn't to slag on consul. I think consul is like chocolate and peanut butter.
I use it in my mesos cluster and just am genuinely curious how large it can scale!
Started with Serf plus customizations and then migrated over to Consul. Happy to share technical details or answer questions.
but half year ago and no community reaction.
https://news.ycombinator.com/item?id=7604787
https://news.ycombinator.com/item?id=7682173
So by HN standards (https://news.ycombinator.com/newsfaq.html) this is unequivocally a dupe. But the community interest here seems genuine, so we won't apply the full penalty.
I was looking it over and found how leader election is handled in general fascinating. Leader election can be tricky if you have a lot of machines. This example sets a watch on a single 'key'. But this would mean that all followers would be triggered upon lock-release, meaning all your machines will pound consul for a new lock request / leader election fight. To mitigate this, it seems they implemented 'lock-delay'.
In ZooKeeper what you typically do is create a linked-list of watches so that only the follower directly behind the leader gets triggered and so avoid 'disturbing the herd'.
In any case, consul seems great, and will probably use it!