Automated Testing for League of Legends
engineering.riotgames.com
engineering.riotgames.com
The specific thing I referred to as a myth, though? That is completely fabricated, the guy lied in his research. Total myth.
The result would be a dimino-compliant way of establishing that smoking is not bad for you. Or at least, that "smoking is bad for you" is a myth.
Of course, for smoking we have dozens of studies and a strong understanding of the mechanisms, so this trick wouldn't fool anyone. But, just be careful about letting fraudsters determine your opinion on a topic, whether for or against, whether they're caught or not.
As for bugs specifically, there are huge benefits to being able to reason about systems as if they are bug-free. It's a PITA when your libraries, kernel, compiler, upstream API, etc. don't behave as expected. As professional software developers, we can work around them; users can sometimes get really confused when things don't work as expected. Sometimes they're scared to tell you.
Smaller companies can probably bear more bugs, since they can give individual users more attention. Larger ones tend to need to work more reliably.
I'll leave you with a recommendation to read The Startup Owner's Handbook, and listen to the Stanford lectures given by Sam Altman and friends, who talk a lot more about this.
You clearly have never worked in fintech, banks or payments. I personally witnessed the moment we catched a race condition which costed the company n-k euros in missed transactions.
And this isn't even considering mission critical software where bugs can (literally) kill.
(I say "identified" because you don't always fix identified bugs. But you really, really want to have as much as possible identified before production. Production is a terrible place to find bugs for the first time. Yeah, every once in a while an identified bug might have a far worse than expected impact, I've had that happen, but at least in my experience the completely unidentified bugs are what kill you the hardest.)
If you're building rocket engines that carry people, you think like this. If you're building a social networking website, or a food ordering website, or a home sharing website, the time it takes to go from "oh there's a bug" to "that bug is fixed in production" is a matter of hours, if not minutes.
If you've only ever encountered bugs that can be fixed in "minutes", you're either very lucky, or not working on anything all that important. Or possibly, not very perceptive and you've actually got a mess on your hands and you haven't realized it yet. I've inherited systems run by such people; they think everything's hunky dory but it turns out you can hardly figure out how to connect two records in the database together correctly because they were too smart for academic bullshit like referential integrity, and, lo, their database was low on the integrity. Tends to work, except the amount of elbow grease required increases without bound until it exceeds the capabilities of the developer in question, and suddenly, one day, they wake up and their job is in serious jeopardy because of what's going down and what they can't get back up.
You assert that this is the "older" style of thinking, which shows a lack of understanding of your computer programming history. What you advocate is "cowboy" thinking, and it predates what I'm discussing by quite a bit. The entire 1970s was basically run this way, and it shows in those handful of remaining technologies that are still with us, and their tendency to just pound through problems as if errors are an admission of failure.
Netflix runs their Simian Army in production, for example. They sometimes pull infrastructure pieces offline intentionally to test their resilience. A good production infrastructure is incredibly resilient.
Engineering will naturally tend to geek out on infrastructure projects like this and they need to be kept focused on the business case. There has to be some push-back.
What is the cost of even one engineer working full-time on build verification?
The thing you really want to avoid is getting blind-sided by devastating bugs or systemic process problems. So it's a balance.
I've just seen many cases where projects got bogged down as developers built their super-uber-build-test framework. It's easy for people to push these projects through without rationally investigating the cost / benefit, because ... because ... you're not seriously suggesting that we shouldn't test more, you monster!?
These sort of downtimes hurt customer satisfaction and the bottom line. If serious issues arise often enough, customers lose confidence and patience and may leave your product for a competitor. And while you're offline trying desperately to rush a fix, you're losing revenue.
If the service/product you're offering has competitors, quality matters. Sometimes it's not even the measurable quality, but the perception your customers have. Find me a business owner who thinks major bugs in production are not "a huge problem".
"Move fast and break things" - Mark Zuckerberg.
>> We used to have this famous mantra... and the idea here is that as developers, moving quickly is so important that we were even willing to tolerate a few bugs in order to do it. What we realized over time is that it wasn't helping us to move faster because we had to slow down to fix these bugs and it wasn't improving our speed.
They have a test client in which a small group of players can test out new patches and find bugs. Recently when patch 6.87 was released, they practically skipped the test client and deployed it straight to the main client. The result was that there were all sorts of bugs in the game, but they worked really quickly to fix them so that after a few days, they were almost all gone. I think Valve's mentality is that it is more efficient for their large player base to beta test for them instead.
Also, there's been a few times when the game is literally unplayable for one or two hours because Valve pushed a random update.
(LoL is structured around roughly yearly 'seasons' for competitive play and for ranked play rewards. This season they have released a far greater amount of game mechanic changes, which make it difficult to plan competitive strategy and difficult for regular players to keep up.)
The amount you'd need to abstract to make it reusable for other developers would make it nearly useless.
- Staging area for their tests, so many times I've lost confidence in our test suite because of a flaky test. We then delete that test forever. Adding tests to a staging area to very that they are stable is a great compromise
- each test is a class, with 'setup', 'execute' 'verify' - makes the test a lot easier to read and refactor
> In Wood 5 we don't use wards anyway, so I see no problem with this critical failure
100% agreed. There's no point in changing to lens because nobody buys a ward!
Update: In particular this is weird:
> "Tests make use of remote procedure call (RPC) endpoints exposed on the client and the server in order to issue commands and monitor game state. For the most part, tests consist of a fairly linear set of instructions and queries—existing tests cover everything from champion abilities to vision rules to the expected rewards for a minion kill. "
They test "rules for expected rewards" with an out-of-process python test program that connects to the client and server via RPC. Seems unnecessary complex way to test specific game rules.
Is this a symptom of a separate QA team and no developer-written unit-tests?
What Riot is using is a form of integration test where manual user input is emulated to produce a result data set. What you aren't seeing in the test code is "everything else" that was needed to set up a running game state. This technique makes it easy for QA to jump in and see what's happening visually when a failure occurs, eliminating the need to slave away at a checklist.
Additionally, you want to use production server/client code as much has possible, while getting through the tests as quickly as possible (skip front end flow, matchmaking, etc.).
Using client/server endpoint calls are great because it allows these integration tests to be a list of instructions (the same instructions used by production code) to create a very deep test suite.
Integration testing is critical for games, where you have so many independent units interacting and changing each other's states.
Definitely. They're at 130 champions right now. Each with four spell cast abilities, as well as a passive ability, and unique auto attack mechanics on some of them. There are indeed many interactions with specific abilities between champions. Add in interaction with the map and terrain itself, like unit collision and NPC minions and monsters. Then deal with the 154 items (the current count on the main map) players can add to their champion, many of which also interact with abilities, and some of which introduce additional abilities.
It must be both a nightmare and yet very interesting to handle creating a new champion for the game. Once they get past the step of even deciding how the new champion's abilities should interact with other champions, they then need to code it all and manage to verify that everything works according to expectations. Every once in awhile, there are still bugs with how one ability interacts with another. You can plan it all out, but it has got to be easy to miss something. Too many interactions! :)
But if you have a farm that tests v1 of your new champs abilities against all others, you know exactly where the 20% of problems will lie. This is where designers earn their paycheck- two abilities that won't resolve by design, and where the programmer puts 80% of their work.
80/20 rule scales with the product, I suppose :)
I've played this game for 7+ years now and as a programmer love to hear about all this. A Rioter said a couple years ago that they used Erlang[0] and as a longtime Erlang (and now Elixir) admirer and tinkerer that was great to know my favorite game uses it.
[0]https://erlangcentral.org/scaling-league-of-legends-chat-to-...
The APIs that they have for testing alone are pretty incredible; I'd love to be able to play in a mode where you could change level on-demand, create dummy enemies, etc.
It's obviously a stretch as that guy's username is the name of a DotA hero, but hey, rose colored glasses are fun sometimes.
Lots of bugs in Dota 2 seems to be happening because of manual testing.
DotA is balanced with creativity in mind. When an exploitative strategy is found the game is not changed to remove it but instead through some combination of tweaks and modifications, often to seemingly unrelated aspects it's made into something that's only situationally viable.
The reason the game is so deep/interesting is because this philosophy has been in place for many (10+) years now resulting in a game that's much more complex than pretty much anything else out there. DotA 2 retains edge case behavior that was present in the original Warcraft 3 mod. Things that were once engine limitations are now part essential parts of gameplay. Preserving these weird interactions instead of attempting to make the game more regular/comprehensible is one of the things that allows such a high level of creativity in competitive play. There's always a way you can outplay your opponent that's clever. In go parlance: there's always new tesuji to be found.
All that being said I think Riot's efforts here are quite laudable and DotA could certainly find benefit from more rigorous testing.
None of your comments so far have contributed anything to this discussion. Please try to keep to the spirit of HN when commenting.
Verifies:
- KogMaw deals less damage to non-lane minions
- KogMaw deals percentile magic damage
- KogMaw deals normal damage to lane minions
and the verify method has three assertions. In my opinion, this should be three tests, each with only one assertion.I bet the tests aren't pixel-by-pixel frame-by-frame reproducible (they are talking about a one week staging period for tests to prove themselves, so surely there's some non-deterministic wiggle room in remote controlling the full graphical client), so you really want to cover all assertions together for each run.
Making each an individual test would triple the test time, pollute the test codebase and not bring any advantage.
Ideally, if just one assert fails, it will produce an error message/logs with enough info to tell what went wrong. Having tests produce great error messages is way more valuable than arbitrary rules about assertion density.
Would it really? None of what you said here is true.
It's just hard sometimes to set and follow standards when you're pushing to demonstrate value. Not impossible, but hard.