10M Concurrent Websockets
goroutines.com
goroutines.com
For comparison, Chris McCord (the creator of the Phoenix framework) did some work on getting 2M concurrent connections with real-world(ish) data being broadcast across all nodes: http://www.phoenixframework.org/blog/the-road-to-2-million-w...
http://www.phoenixframework.org/blog/the-road-to-2-million-w...
If the kernel is bypassed, I have reason to believe Go can handle the line rate of 16.7M packets/sec(on 10gbps nic.) I would love to integrate Intel's DPDK with Go. I've done some ground work on this before - it's very feasible to just replace the Go network stack in a transparent way. If any company is interested in sponsoring that as an open-source project, my contact info is in my profile.
Do you plan to write the TCP stack from scratch? If not, are you planning to leverage one of the golang unikernel back-end ports?
10Gb/s (even per-core) doesn't seem like much of a selling point for a network stack on it's own, given where the rest of the industry is.
I'm actually strongly tempted to just to glue Seastar and Go together. That's the lowest risk approach.
Good luck with your work - I'm curious to see the result.
Sending a ping every 5 minutes at 10M connections was roughly the limit of what the server could handle. This ends up being only about 30k pings per second, which is not a terribly high number.
12 mil sockets, 200k pings / sec, 12 core server 54GB mem use, 57% cpu
https://mrotaru.wordpress.com/2013/06/20/12-million-concurre...
In that benchmark, pings are actually 512-byte messages. So, MigratoryData published 200,000 messages/sec to 12 million clients at a total bandwidth of over 1 Gbps.
BTW that benchmark result has been improved recently in terms of latency. The JVM Garbage Collection pauses were almost eliminated using Zing JVM from Azul Systems and the MigratoryData server was able to publish almost 1 Gbps to 10 million concurrent clients with an average end-to-end latency of 15 milliseconds, the 99th percentile of 24 milliseconds, and the maximum latency of 126 milliseconds (computed on almost 2 billion messages):
Mihai
Zing JVM is not a requirement for the MigratoryData server. MigratoryData includes by default Oracle JVM, which is currently used in production. With Oracle JVM we currently obtain quite decent Garbage Collection (GC) pauses for Internet applications, so the latency is not too much impacted by the GC pauses which occur from time to time.
Zing JVM could be used in certain projects to reduce near to zero those GC pauses, so the latency is not impacted even in the worst case, when a GC occurs. So, if Zing JVM is necessary for such a ultra low latency application, Zing JVM should be licensed separately.
For example my R710 shows 16 threads while it's only 2x4 real cores.
While the 10M figure is impressive, this doesn't sound practical or useful.
- 10 million concurrent connections
- 10 gigabits/second
- 10 million packets/second
- 10 microsecond latency
- 10 microsecond jitter
- 1 million connections/second
However it was for raw connections, not for websockets.
Hetzner, a German hosting company, have some really good root servers, for about 120€/month you get a beast of a machine with 16 cores and 128gb of RAM. If you add some basic load balancing, you can achieve 10M for a really low price. Add something like Docker and you can get a PaaS-like setup, that handles millions of connections without breaking a sweat.
Best example I know of squeezing as much as possible out of the bare metal servers would be StackOverflow, which runs of something in the ballpark of 20 servers(excluding replicas) IIRC.
Usually by bypassing the docker networking and using Host only network, but then you lose a lot of the benefits of containers in the first place.
For example, weave or calico networking layer on top— which add a fair bit of latency if your aiming for 10M connections— makes scaling containers quite easy
I would imagine that if you plan on using Docker in your infrastructure seriously, you are aiming for a multi host setup with many containers spread throughout— and can settle for 100k connections per container easily.
If we're talking about these kind of specs, something tells me Erlang would do even better.
I'd want it to be something with a more safety guarantees -- could be compile time (strong type system, Haskell, Rust) or strong runtime fault tollerance -- Erlang/Elixir and other languages on the BEAM VM.
If you're going to use a machine with 208 GB of RAM you may as well buy a better network card. All major vendors of low latency adapters (which in this case means handling lots of small messages) have transparent drivers for Linux kernel bypass now. For example see OpenOnload. Then you avoid the kernel with no code changes.
Once you remove the kernel from the critical path, you should be concerned with I/O resources, which is already illustrated by the article's note about how they get 5x better performance by using four machines with the same total core count. What that does is buy you four times the I/O resources, like memory channels and PCIe lanes.
Not even close. Layer in X509 authentication and flaky mobile networks and now you're into tuning backoff while dealing with the CPU overhead of the TLS negotiation.
Also of note, AWS' network use to flake out at around 400K TCP connections from the public network. May no longer be the case but it certainly was ~2012.
If you have these requirements (no read, only write, zero connection state) you would be best served by doing userspace raw IO with something like DPDK instead of the huge machinery that the kernel spins up for every TCP connection.
[1] https://developers.google.com/web/updates/2015/03/push-notif... [2] https://www.raywenderlich.com/32960/apple-push-notification-... [3] http://caniuse.com/#feat=push-api
The only interesting thing here is that with that 10M connections figure, each connection handler goroutine contained at least 2 channels (sub and t), and they seemed to scale (no mention of latency though).
t := time.NewTicker(pingPeriod)
Maybe if you hold all of the connections in one array and create one timer for the ping, things would be faster.
Curious how this compares to netty...
yeah not exactly difficult
Now git off of my lawn, I need to take a nap.
However, Go's ability to do coroutines + multicore certainly makes it easier to take advantage of the system's full potential, and this ability is not terribly common in other languages (at least not in such an automatic way). Of course the same result could be achieved with any event-driven language in a less automatic way.
Arguably any functional language does this even more automatically than Go does.
The language doesn't matter, as long as its compiler can expect purely functional code and the compiler is written to take advantage of that code.
Rust kind of succeeds in having it both ways -- the programmer decides what runs in parallel, but the compiler won't allow race conditions.
I can't imagine any better scenario than just writing code that can be automatically parallelized, though.
I'm not holding out much hope for automatic parallelization to save us, or even help all that much. The only remaining hope is that a language written from the beginning to afford more automatically-parallel programs will help, but those efforts (Fortress most notably) also have not gone well. https://en.wikipedia.org/wiki/Fortress_%28programming_langua... , http://web.cs.ucla.edu/~palsberg/course/cs239/F07/slides/for... starting page 33
It's a pity, really. But unfortunately, I wouldn't be promising anybody anything about automatic parallelization, because even in lab conditions nobody's gotten it to work very well, so far as I know. It's a weird case, because my intuition, like most other people's, agrees that there ought to be a lot of opportunity for it, but the evidence is contradicting that.
Anyone who's got a solid link to contradictory evidence, I'd love to hear about it, but my impression is that research has mostly stopped on this topic, because that's how poorly it has gone.
It is, however, easier to use, but the context switching and stack overhead would be too great at anything close to 10M connections.
I'm very much a fan of this sort of work--it's obviously useful for things like data collection from sensors--but I can't help but wonder if it has any practical business value for early-to-mid stage startups.
We sort of take it for granted that you will have to rebuild your architecture once you hit that point. What you gain for that is a platform where you can be productive and iterate quickly.
What the Phoenix example shows is that without a huge loss of productivity up front, you can put off the rewrite much longer.
Presumably, Go would provide similar benefits.
When I was first playing with multi node Phoenix stuff I was amazed. It was too easy to scale horizontally.
I'm not sure what the distribution story is for Go...
apt-get install -y redis-server
echo "bind *" >> /etc/redis/redis.conf
/shiver....And your new redis server joins all of the other unprotected ones directly connected to the net with no protection or real security.
M:N schedulers and the primitives they use for their 'lightweight' threading are always going to use more memory than a single-threaded event loop.
These should give you some idea on what is going on there:
https://golang.org/src/net/fd_mutex.go https://golang.org/src/net/fd_unix.go#L237 https://golang.org/src/syscall/exec_unix.go#L17
Also the FD mutex that need to be taken before each read and write is nasty (is it just to prevent a race with close or is it for something else?), but at least that won't usually require any syscall and on sane applications could be optimised with a Java-like biased lock.
It's using green threads as well, with 1 "process" per connection, right? Is it not the same kind of scheduling?
I don't understand what the grandparent is saying though. Yeah, if you create a thread for every request, then you're going to be killed by the memory overhead. This isn't unique to Go, and I remember it being a problem when many noob Java programmers would spawn a thread per request before the introduction of java.nio.
The downside with Twisted, Tornado, or whatever in Python is your code isn't parallelized. It's concurrent, yes, but you aren't taking advantage of your multiple cores without forking another Python process due to GIL.
Go, Scala, Java etc. are truly multi-threaded and compile to native code from commandline or via JIT. Saying your performance advantage is due to switching from Go to Python is a spurious claim. You weren't doing it right.
Sequential code tends to be more maintainable, readable, etc. vs. evented code.
So while you are right about evented systems, this isn't pointless.
This is simply not true, I'm sorry.
As for pointless - the whole thing shows how Go is useless at 10M connections and nothing more.