K6: Like unit testing, for performance
github.com
github.com
What I'm really excited to have discovered is the `HAR` format.
https://en.wikipedia.org/wiki/HAR_(file_format)
Over the past 10 years I have glued together project-specific variants of "parse some proxy or WAF log and spit out test or automation data" at least 4 or 5 times. Including as recently as a couple of weeks ago:
https://git.sr.ht/~tuxpup/parse-charles-log
Knowing HAR files existed would have saved me quite a bit of work and left me time to improve my automation or testing. I'm very happy to have discovered this.
You can use it to archive websites and recreate them offline, Internet Archive supports HAR files for archiving.
You can use them to recreate user scenarios for acceptance testing and performance tests, via tools like K6 (K6 which also supports InfluxDB, so you can gather the metrics yourself and have your own UI [via Grafana or anything else that supports InfluxDB])
You can use them for troubleshooting all sorts of user/client issues if you have a easy way of gathering the HAR files, or if your users are developers.
They're simply a invaluable tool in web development these days, if you deal with web frontends/backends.
> JavaScript is not generally well suited for high performance. To achieve maximum performance, the tool itself is written in Go, embedding a JavaScript runtime allowing for easy test scripting.
How is it possible that pure go JavaScript interpreter (goja) with bindings for net/http and some reports would be faster than the same tool written in nodejs using its http-client (which if I remember correctly is written in C)?
I don’t mean to downplay the importance or usefulness of k6, I just find their reasoning behind choosing go somewhat contrived
You can find some basic performance and other comparisons between load testing tools in this very long article of ours: https://k6.io/blog/comparing-best-open-source-load-testing-t...
And some advice for squeezing the maximum performance out of k6 in here: https://k6.io/docs/testing-guides/running-large-tests
Anyway, what I've seen when comparing the performance of tools, is that Artillery, which is running on NodeJS, is perhaps the wors performer of all the tools I've tested. I don't know if it's because of NodeJS or that Artillery in itself isn't a very performant piece of software (It also consumes a lot of memory, btw).
If you want the highest performance, there is one tool that runs circles around all others (including k6), and that is wrk - https://github.com/wg/wrk - very cool piece of software although it is lacking in terms of functionality so mostly suitable for simpler load testing like hammering single URLs.
(I don't know how fast wrk2 is, haven't benchmarked it)
IMHO that is about JS-memory usage and GC, which never really get's proper attention.
You can have really fast algorithms, if you keep creating and forgetting millions of objects, which most JS-frameworks do, you will have lags and GC-pauses.
In build scripts, it's important that you can start up each task (k6 instance) very fast.
No, and definitely not when comparing with the time it takes to run load tests.
Is this a general tool, or is it specific to web development? e.g. Would this be an appropriate tool for Blender to detect performance regressions in their rendering?
I like the design of running each virtual user in its own Javascript VM inside of a Go process. I can just install a single binary, but still write tests in a high level language (and not rebuild the load testing framework for every change).
An idea that's been kicking around in my head for a while is extending Go applications by embedding a WebAssembly VM, exposing relevant hooks, and letting users add whatever they want at runtime. That way, you don't have to bundle a bunch of crap into a statically linked binary, users can go add that later :) K6 seems to validate that this is a good idea; while not WebAssembly, it sure works well. (I would want to write my plugins in Go, not Javascript... but the idea is good.)
We were initially looking for something slightly different though: we were interested to have perhaps less tests, but tests that would run much much more often (like every seconds or couple of seconds), in a continuous manner. Tue goal was to have something at the same time like a healtcheck (is it still working), like a performance test (does it answer in a timely manner) and like a validation test (does it answer the right result - the endpoints we wanted to test do "complex calculations"). Our best answer so far was to wrap K6 in an infinite loop, but I wonder if there could be something smarter.
> tests that would run much much more often (like every seconds or couple of seconds), in a continuous manner.
You can do that, just use an arrival-rate executor that runs an iteration every second, with a test duration of 365 days or something like that :) See https://k6.io/docs/using-k6/scenarios/arrival-rate
> Tue goal was to have something at the same time like a healtcheck (is it still working), like a performance test (does it answer in a timely manner) and like a validation test (does it answer the right result - the endpoints we wanted to test do "complex calculations"). Our best answer so far was to wrap K6 in an infinite loop, but I wonder if there could be something smarter.
You can certainly wrap k6 in an infinite loop. Nothing wrong with that, though you can probably use the `scenarios` feature (with long `duration` values) to achieve it without wrapping k6: https://k6.io/docs/using-k6/scenarios
I use it with influxdb and grafana.
Set it up with one virtual user running for 6 hours, requesting different endpoints, you should get the 1req/sec.
If more complex flows and assertions on responses are required I go for Gatling. So maybe K6 is more comparable to gatling.
Disclaimer: I'm one of the people who created k6
Warning: The article is pretty long...
Personally I use autocannon which in my testing is faster than k6 and easier to use with a node.js environment. It is only a npm package, not a full software you need to install on your machine. Autocannon in my testing is 2 orders of magnitude faster than gatling, but you should do your own testing. PS. I have no affiliations to autocannon, I'm just a very happy user.
The only drawback compared to k6 is lack of websockets, which is unfortunate.
In my testing, Gatling isn't catastrophically slow by any means. Wrk, which is faster than any other tool I've seen, is about one order of magnitude faster than Gatling, in terms of raw RPS generation. I find it very hard to believe a NodeJS-based tool executing JS would be faster than the fastest tool written in C that just hammers static URLs.
If you just run it in your own CI/CD pipeline: no.
If you want to offer this as a SaaS: yes.
If you want to extend this somehow: yes.
If you statically or dynamically link any part of the k6 codebase with some other code (a derivative work), the license’s copyleft virality is triggered and that other code also needs to be made available under AGPLv3. When it comes to using the k6 binary and interacting with it from another process or over a network it is not certain exactly how the copyleft virality would apply from what I’ve been told by lawyers with OSS license knowledge.
That said, I do have a couple questions:
- Would a non-AGPLv3 open-source or source-available project using K6 be incompatible with AGPLv3 given they might "distribute" CI/CD along with their open-source or source-available code?
- If a closed-source project publishes results from K6, are they now in violation of AGPLv3?
> Thanks. I must say, I prefer weak copylefts for my projects (MPLv2 with the incompatibility clause), so my concern isn't copyleft per se, but the virility of it.
I need to read up some more on MPLv2, but from what I've read now it looks like a license that could fit k6 as well.
> - Would a non-AGPLv3 open-source or source-available project using K6 be incompatible with AGPLv3 given they might "distribute" CI/CD along with their open-source or source-available code?
No, using k6 as a tool in CI/CD would not be a violation of AGPLv3 or trigger the virality of the license in regards to the non-AGPLv3 open source or source-available code base. Any k6 test scripts you write would also not need to be licensed under AGPLv3. Gitlab built an integration with k6 which could be seen as a data point in agreement with this view: https://docs.gitlab.com/ee/user/project/merge_requests/load_...
> - If a closed-source project publishes results from K6, are they now in violation of AGPLv3?
If we're talking about publishing results from testing of the closed-source project, then no, that would not be in violation. There's however an undefined gray area around clause 13 "Remote Network Interaction" in AGPLv3 in terms of what is allowed when for example offering a SaaS product based on an AGPLv3 licensed component such as k6. Clause 13 has not been tried in court afaik. It could seem that clause 13 would provide some protection for a business like ours from other companies building commercial solutions on top of k6, but from what I've been told by lawyers we shouldn't count on that, so we're not (anymore).
We have had many internal discussions on this topic, as well as with other companies and individuals in the open source space; whether to go the route of a source-available license, effectively restricting commercialization possibilities of k6 by other companies, or going the open core route and restrict what we release as open source. Everytime we've had this discussion internally we've come back to open source being the right choice for us, given the type of product we build and where we think we can capture value.
We'll look into MPLv2.
MPLv2 and other related licenses such as Eclipse Public License v2, Erlang Public License, do have the advantage of being well understood and in some cases auto-approved for use at various enterprises and thus a good midway between MIT / Apache and GPLv3.
That said, you'd be right to lean more-copyleft (going Server Side Public License, for example) if K6 is a key product (and not a complementary product), though Bryan Cantrill thinks you might be better off closing up the source in that case: http://dtrace.org/blogs/bmc/2018/12/14/open-source-confronts...
That's unfortunately not the case for AGPL. There's a reason a lot of companies have policies banning any AGPL dependencies outright, Google probably being the most prominent example: https://opensource.google/docs/using/agpl-policy
AGPL is designed to be extremely viral and has not been tested in court. The definitions of boundaries between your code and AGPL-licensed code are not well understood. Consider that MongoDB, probably the most popular AGPL-licensed project in use before 2018 when they switched to SSPL, had to explicitly publish their drivers under Apache because communicating with an AGPL dependency over a network was not sufficiently distant enough to prevent infection. https://www.mongodb.com/blog/post/the-agpl
As engineers we need to be aware of licensing implications of code we depend on. I am obviously not a lawyer, and obviously your employer's legal team are the people you should be talking to if you have an absolutely critical dependency on an AGPL-licensed project.
As I mentioned in reply to another comment, from a first look at MPLv2 it seems like it could be a good fit for k6 as well.
You said: > We chose MPLv2 specifically to address licensing concerns.
Is that something you've encountered while building Artillery? We've not met a lot of pushback from legal teams at companies using k6.
You can export HAR files from all web browsers and a lot of proxies and other tools. Then you can convert that recording to a k6 script with our HAR converter: https://k6.io/docs/test-authoring/recording-a-session/har-co...
We also have browser plugins to directly record browser traffic and import it as a k6 script in our cloud service: https://k6.io/docs/test-authoring/recording-a-session/browse...
However, you don't need a cloud subscription to run the tests, you can copy-paste the generated script code and `k6 run` it locally, without paying us anything.
Regarding debugging, that's unfortunately unlikely to come any time soon... We'd need support for that in the JS runtime we're using (https://godoc.org/github.com/dop251/goja) and then we'd need to figure out how to expose in in k6, while dealing with potentially hundreds of concurrent JS runtimes (VUs). Not impossible, just unlikely to land anytime soon...
Can you DIY a solution that will let you instantly scale up from running tests locally to running them on 500 workers in any of 13 different geographical regions? With no servers to manage or maintain whatsoever. Sure you can, but it's not the best use of time for a lot of teams.
See the example in https://github.com/loadimpact/k6/#checks-and-thresholds. If, at the end of the test run, some of these rules (encoded as `thresholds`) were unsatisfied, k6 will exit with a non-zero exit code (and thus fail your CI check): - The 95th percentile of all HTTP request durations should be be less than 500ms - The 99th percentile of all HTTP request durations that were tagged with `staticAsset:yes` should be less than 250ms - The failure rate of the checks should be less than 1%, although if it's more than 5% the test will abort immediately.
Gil Tene (Azul Systems) has argued convincingly [1] (slides [2]) that monitoring tools get latency measurement wrong because sudden spikes aren't represented correctly in timings because of averaging and the wrong use of percentiles.
He argues that percentiles simply aren't useful, because, statistically, most requests will experience >= 99.99-percentile response times. All percentiles lie; the only truly useful and realistic number is, in fact, the 100th percentile.
He also argues that the most revealing way to show latency is with a complete histogram.
I haven't used K6 yet, but I've been looking for a good load-testing tool, and intend to try it out.
[1] https://www.youtube.com/watch?v=lJ8ydIuPFeU
[2] https://www.azul.com/files/HowNotToMeasureLatency_LLSummit_N...
It has been a while since I watched these Gil Tene talks, so I might be missing something, but I think the only remaining task we have is to adopt something like his HDR histogram library [2]. And that's mostly for performance reasons, though we'll probably play around with the correction logic as well.
k6 allows the user to choose[1] which metrics are relevant for the particular test. By default, it displays max or p(100), p(95), p(90), min, med, and avg. User can specify other values such as p(99.995)
It's also possible to create completely custom metrics[2] to track whatever is relevant to the user.
k6 allows the user to change almost all aspects of execution, tracking, and reporting.
This sounds more like integration testing than unit testing. Something like Google's micro-benchmarking library [1] is more like unit testing.
We have a more comprehensive review of open source load testing tools in our blog (including Tsung): https://k6.io/blog/comparing-best-open-source-load-testing-t...
Specifically, and with the disclaimer that I haven't used Tsung so much, I'd say the biggest difference is that k6 tests are scripted in a real language (JS) as compared to Tsung's XML.
But k6 also has a modern CLI UX that beats Tsung (and most other tools) by a mile, more integrations now and possibly support for more protocols, or at least more relevant ones (not sure if Tsung supports HTTP/2 for instance). Plus k6 is being very actively developed.
From what I've seen of Tsung I really like it though. Seems like a very performant and solid piece of software. But like mentioned earlier, it was made in another era. Just like Jmeter. If I ranked tools in terms of ascending "DevOps:ness" and UX for developers it'd probably be something like:
Lowest: Jmeter, Tsung
Low: Gatling
Medium: Apachebench, Siege, Hey, Wrk (simple testing)
High: Artillery, Locust
Highest: k6, Vegeta (though Vegeta is only for simple/static tests)
I wrote an article about this topic also: how most load testing tools we have are just not geared for developers and DevOps: https://k6.io/blog/another-load-testing-tool
On the `k6 cloud` side, we have executed 500k+ VUs. The most RPS I achieved with k6 was 4 791 928 (~4.8 million requests per second). That test lasted for 6 min and generated 1.5 billion requests in total.
More is possible, but we didn't push further.
Source: I'm one of the guys behind k6.
Local execution with cloud output is also possible (k6 run -o cloud), but that's not free.
To give you some idea of what k6 cloud can do, here are screenshots of a few large scale tests https://imgur.com/a/NtXsc1a
1. 500 000 VU test run with 1.5B requests
2. 12h soak test run with 5k VUs
3. 48h soak test, with 5k VUs
4. 20k VU local test (with -o cloud)
Adding MQTT should be rather straightforward.