Wi Flag (2002)
asheron.fandom.com
asheron.fandom.com
I felt more and more that dice values of specifically 1 and 6 were harder to come by than other values, so one day I sat down for a few minutes and logged the value of 100 or so dice throws. Turns out I was right, the distribution was not uniform: 2-5 were fine, but it seemed that it was twice as hard to get a 1 or a 6 compared to any other value.
I even did a chi-squared hypothesis test, because I was crazy. (And because I was studying for my statistics minor in university at the time).
Seeing the result, the problem was pretty clear, without even knowing the source code. Almost certainly, they had a random number generator giving a number in a uniform range, let's say from 0 to 1, and did something like the following to get a dice value from 1 to 6:
value = round(random()*5)+1
Do you spot the issue?The problem is that round() rounds to the nearest integer. So any value between, say, 1.5 and 2.5 (exclusive at one end) rounds to 2, 2.5-3.5 rounds to 3, and so on. Add 1, and you have the dice value.
But the ranges for dice values 1 and 6, are only 0 to 0.5 and 4.5-5 respectively, so ranges that are only half the size.
The fix should be extremely simple:
value = floor(random()*6)+1
The crucial point being to use floor() or ceil() instead of round(), to only round towards one direction.I wrote up exactly this in a short email to the company that made the Yahtzee game. They did not reply and just silently fixed the bug, right after my email. I was disappointed and stopped playing.
I love this. It's one of those things that's entirely intelligible to me but I imagine would be next to impossible to explain to a Martian.
But here's another layer, though. What if it isn't?
Our science fiction is filled with this fairly human-chauvinistic worldview where even in a hypothetical future where we are rubbing shoulders with many different alien species, many of them technologically superior to ours, some of them still admire or begrudgingly respect us for our ingenuity, spirit, some je ne sais quoi.
That we're illogical and inconsistent and jealous and petty and insecure, but also humble and funny and empathetic and creative.
But...we have no evidence that this would be unique about humans whether sentient intelligence is common throughout the universe or rare. In fact, it's quite possible, that this is a necessary prerequisite, side effect, or consequence of conscious intelligence.
So this hypothetical martian maybe would just smile, nod and say "Ah yes, the company was embarrassed, or didn't want to risk liability, or didn't want to risk their reputation being compromised. But OP just wanted a thank you. He wouldn't have wanted money, just a 'hey thanks for finding that for us'. And it soured him on the whole experience. Been there, my new human bro, been there."
edit: I wrote this before I saw the rest of the replys, and HOLY, there's a bunch of humans that don't understand why OP stopped playing.
I can’t wait for AI psychologists. AIs hired to make AIs talk about their reasonings because they are stuck in feedback loops. AIs having to be debugged. AIs being too selfish or too generous or too insecure that they keep deflating what they try to say despite being of extremely high value. AIs suddenly cracking nonsense when we need them the most under high pressure. I’m persuaded all these are required aspects of general organic intelligence.
One could argue it was a coincidence, but I still like to believe that I influenced a major site's uptime. Mostly because they told me I crashed the site and asked me not to play any more jokes like that again.
It is not the players fault patches take time to get through QA. Nor should there be an implicit valuation curve to free bug reports from players.
Or their internal communication was haphazard and not designed to handle this case. I could totally imagine the bug report getting passed informally through a game of telephone and losing the connection to the external reporter (e.g. at some point the report got paraphrased, so the developer who fixed it doesn't know where the report came from, and the person who received the initial report doesn't know if it had merit or if it was fixed).
Worst case you thank someone for something useless.
They included the fix in the next version of the tool. I later noticed it was incomplete, but this time I just made the correction locally.
Technology Connections
The relevant part about the old incandescent bulbs is at 3mins 14secs
1. taking a random integer between 0 and 2^31 − 2,
2. dividing the integer by 2^31 − 1 to get a double-precision floating-point between 0 and 1,
3. multiplying this floating-point with b to get a number between 0 and b,
4. taking the integer part of this new number to get an integer between 0 and b − 1.
Do this for b = 2^31 − 1, and 50.34% of outputs are odd numbers, where you'd expect (almost) as many evens as odds.
Ended up doing a bit of a write-up here: https://fuglede.dk/en/blog/bias-in-net-rng/
Another fun thing about their RNG was that it was apparently supposed to follow Knuth's provably useful additive PRNG which, given n − 1 randomly generated numbers, generates a new number by taking the sum of the (n − 24)'th and (n − 55)'th numbers modulo some specific number; back when I looked, Microsoft's implementation for some reason uses 34 instead of 24, probably just a typo, with the unfortunate side effect that Knuth's theoretical guarantees are out of the window.
As soon as I saw that, I knew exactly what the problem was. I'm glad I was reading on a small mobile screen, so I didn't feel like I was cheating by getting a hint from the rest of your comment.
Yes, voice of painful experience here!
> The crucial point being to use floor() or ceil() instead of round(), to only round towards one direction. [emphasis added]
Not quite. If the random() function in your language returns a value in a half-open interval like Math.random() in JavaScript or random.random() in Python, only use floor(), never ceil().
Those random() functions return a value n where 0 <= n < 1. In other words, n may be 0 but will never be 1. So floor() is always what you want for a random function like that.
A somewhat related tip for anyone who implements a progress message like "nn% done". You should always use floor() for that number as well. Quite often I see a case where someone has used round() instead, or possibly ceil(). And then what happens is your notification message says "100% done" when the process is not 100% done. It makes me a little crazy when I see that. Use floor() and you won't have this problem.
The video game Destiny has this problem with how it displays objective percentages - it will round up, and there's a quest every now and then that shows progress as a percentage and has a high enough denominator to trigger this.
Wi came to one of the player gatherings with little printed out cards and would hand them to people and say, "You've been Wi flagged!"
The fondest of memories.
I remember people theorizing it was due to a short name (only 2 characters), and so would create longer usernames to try and avoid the "curse"
In retrospect, I give a lot of credit to the fact that we were young and dumb and didn't know any better. I've been revisiting a lot of the stories from back then and so many of them end up with us saying, "I dunno, let's see what happens!" and not being dissuaded by "best practices" or even common sense.
Also lots of credit goes to the early internet era when people were a LOT more forgiving of, well, everything.
Countless memories I could keep going on and on, but what an experience!
It's amazing to me how many people are still friends with their patrons/vassals from 25 years ago.
I also really enjoyed the periodic story events that had really dramatic impacts on the world, like the shadow invasion. It was a great game, especially for its time. So thanks for whatever part you played in its creation.
I was on the design team, so was directly responsible for a lot of the shadow invasion stuff (if you ever saw the big bad Bael'Zharon running around in the live events, that was me!) and other patches for the first 2 years of its lifespan.
Weirdly, I work on WoW now with my career having come full circle after having not worked on MMOs since the mid 2000s. :)
I still use AC and to some extent AC2 as an example of how wild, weird, dynamic and interesting MMOs were before the EQ formula won through WoW's success.
Definitely miss the "weird" MMOs of the early 00s...
A theme of a bunch of the comments is that the internet / audience makes this sort of thing impossible these days. One of my white whales for game design is figuring out if mystery, especially in multiplayer games, is still possible in a meaningful way.
You remove the “meta” from the game because every character has a different meta that is just a little too cumbersome to figure out in any way other than “feel” it.
I think growing up during those years of transition, right before the Internet became mainstream and ubiquitous, was a huge boon. Sure, the early MMOs were far from the first international social forum enabled by the Internet, but they were right at the technological frontier at the time. There was something special about inhabiting this massive, 3D virtual space alongside people from across the world, and having that experience be just as novel to everyone else as it was to me.
You couldn't replicate that today, and growing up with the world at your fingertips on a pane of glass as a taken-for-granted fact of life must be a very different experience.
I did a bunch of writing, news posting, collecting information for the monthly patch summaries, and then they figured out I could code and I ended up building tools and maintaining various bits of the site. I think I ended up owning the item database code for a while? I remember hearing that one got used at Turbine, because it was superior to what you had internally!
I’m now married to Kelly Heckman (Ophelea), who was site manager on CoD for a while, and my first real job was at a social gaming startup, getting in the door with the help of her network. That set my career on the path it is now, so you can draw a direct line from picking up that game box to where I am now. So yeah, thank you and the rest of the team :).
The smallest world stays small. :)
Lately, I've been getting more and more into esoteric topics and I keep thinking about this games lore... There is so much occult knowledge/history baked into the entire thing and I honestly still use the game as a reference point for some many things.
I'm curious, do you know how the games lore was developed and were there ancient texts that were inspirations for the events?
Thank you!
I was very not-involved with the lore/story (to the point, where I think I wrote a handful of notes and stuff before it being decided that I should just focus on gameplay, and someone else would cover the lore for me!), but there were 3-4 people primarily responsible for it and they were all very much fantasy literate, so it wouldn't surprise me at all if some of the inspiration were from those sources.
And yeah, Ash Gromnies were the bane of so so so many players.
I haven't really given it any serious thought but I'd likely start with trying to more strongly codify the good parts (incentivizing smaller circles inside of the larger structure, making systems for patrons/vassals to play together in more meaningful ways, etc) while highlighting the positive actions that players could do / benefit from. I don't want to say AC was TOO opaque but a lot of it definitely suffered from being over designed for a very hardcore market.
Loved AC. Bots kind of killed it IMO (something I feel bad about contributing to), but also new MMOs (WoW decimated the competition when it came out). I remember a portal / recall trick with two vendors, where one would sell a specific thing cheaper than the other would buy it. This was before eBay cracked down on virtual goods. Interesting times.
Similar with the min-max:ing of the class-less system. :) IIRC, the aforementioned "dagger warrior" was designed around two things: double attacks at max attack power, and mobs not being able to land a hit. The perfect glass cannon that made it possible to survive AC's harsh lands even when under-leveled - or die instantly. :)
The size of the world was another thing that drew me in - another kind of "grind", if you will. Seeing other players whoosh:ing by - literally running by, in a mount-less world - was pretty hilarious, but the fact that it took time to get around was only a good thing IMO. I want a game world where the only option is to carefully navigate a large, dangerous desert to find the missing ingredient. The current theme-park MMO juggernaut seems to lock most things behind some "boss" and be done with it, which makes the game world pretty void of other players in most places. You just move on to the next quest hub and leave old content behind for ever. This also makes the world, regardless of its actual size, only feel as big as the current "zone" (I don't like "zones").
But in the end, the developer needs the game to be profitable and the players want the most out of their money spent. If my wants belong with a minority group the game probably won't cater to me.
Not really sure what my point is (old man yelling at clouds with contempt for "instant gratification" perhaps), but thank you for AC and the time you spent developing it! :)
I think I mentioned this in another comment and certainly a ton over the years - but a lot of the magic of AC was that it was made by a bunch of people who had never made a game before, much less an MMO, and there were very few ingrained lessons, so we were foolhardy enough to just do things the way it felt we should, player behavior or other consequences be damned. It was built on hopes and dreams and naivete and that made it beautiful and flawed.
But also yeah, once something ships to players, it's now "theirs" and not "ours". We stood in pretty stark contrast to EQ's "you're in our world now" philosophy, again, for better or worse.
But that would be something only a crazy person would do.
It's funny you mention Valheim - one of my groups in had a core of people who I met back in the AC days and their friends. We got to some real old man gaming that probably annoyed their friends with all of the "back in our day..." stories while we were running around the landscape, running from trolls, etc.
Any idea how the targeting algorithm was chosen? It does not behave the way this letter says that it should.
AC is one of my favorite games of all time (neck and neck with Ultima II (I'm old)). I still play on the emulators from time to time - endlessly searching for more Hoary Mattekars.
I played on Thistledown and ran a little portal bot named Stip Dickens an Ayan Baqur that helped ferry people into the Shard of the Herald event. That was a lot of fun.
The loot system was also one-of-a-kind. I don't see that level of randomness in many other games.
The game felt like the writers and staff were highly literate and well read. I now know so many real life herbs and plants due to their use in spellcasting. And AC taught me words like "Mnemosyne".
And each month we would get several pages of amazing lore. (Side note, there's an ongoing web serial called "The Wandering Inn" that reminds me quite a bit of AC. People dropped into a video game like world, insect people, references to a "Zeikhal"(similar to a town name)....
So, I just want to thank you for your work on AC. It's had a profoundly positive impact on my life.
But also thank YOU for participating in some of my very formative experiences of my life as well. The players really made it a joy to work on.
Lost so much sleep power levelling on a rock with those ape like creatures and just spamming magic to discover new combinations. The discovery and wonder does not match up with the graphics I see now when I search for screenshots and videos. Beta and early WoW almost had the same effect but your first is always special.
On patch day, the office banter usually centered around the things we were seeing popping up as we'd reload Maggie the Jackcat's community-sourced patch notes throughout the day [0].
Sometimes we'd even download the patch and log in for a few minutes, just to poke around. (Always being careful not to do anything _real_ due to the risks the infamous patch day server rollbacks.) Then there'd the be waiting for the Decal updates.
I still remember that my main was a tank archer build, which was pretty much untouchable going toe to toe with the mobs in PvE as long as the stamina held out. I went through a _lot_ of stam potions, though!
I remember also finding somewhere about how to extract the terrain data from the game. I scraped the coordinates for various destinations from some place and wrote a little OpenGL viewer that would let me do flyovers with the various locations marked with labelled bullets.
Good times!
P.S.: Anyone remember the Drudge Dance [1]?
[0] http://www.thejackcat.com/AC/Culture/Dereth/Dereth.htm
[1] https://www.angelfire.com/rpg2/dragons/Drudge_Dance.wav (direct link doesn't work; copy and paste URL instead)
User IDs were random UUIDs. Let's say we want to release a feature to ~33% of users; we take the first 2 characters of the UUID, giving us 256 possible buckets, then say that everyone in the first 1/3 of that range gets the feature. So, 00XXX...-55XXX... IDs get it, and 56-FF do not. This works fine.
However, if we then release another feature to 10% of users - everyone from 00-1A gets it, 1B-FF do not. That first set now has both features, and 56-FF have none. It turns out you can't draw meaningful conclusions when some users get every new feature and some get none at all.
feature1_enabled = hash(user_id + "-feature1") / HASH_MAX < 0.3; # 30% of users.
feature2_enabled = hash(user_id + "-feature2") / HASH_MAX < 0.1; # 10% of users.But if need that kind of analysis for usability A/B testing, you most be doing something very wrong on a previous step.
00—1A: have feature flags A and B
1B-55: have feature flag A only
56-FF: have no feature flag
So the actual gotcha here is that there is no cohort for "feature flag B only", right?
This setup can actually be desirable if feature B depends on feature A.
Yep, exactly - by "that first set" I meant the 00-1A group, could have been clearer. Whatever the smallest rollout bucket is, that group is guaranteed to have every single feature.
This was quite a while ago, but I think the actual case we noticed this with was several features released to 50% of the userbase - so every single user either had all or none at once (unintentionally)
And that, as you add more tests, user 00 will always get the test treatment for every test. If you're running a lot of tests which introduce experimental features or changes to workflows, user 00 is probably going to find the site a lot more chaotic and hard to understand than user FF, and that will skew your test results.
those poor cursed users. reminiscent of a joke that used to grace Paul Mineiro's blog (machinedlearnings.com) a few years ago:
> An old curse says: "may you always be assigned to a test bucket."
Also reminded me that I backed Camelot Unchained that's still kicking around in development after all these years.
Players claiming to be cursed, and the curse was real!
A very different sort of event that it reminds me of is the old World of Warcraft plague: https://en.wikipedia.org/wiki/Corrupted_Blood_incident
What I find interesting is the substantial curiosity and effort people will put towards solving questions in video games, but when asked to solve problems in the game of life that we are embedded in, people often seem to have opposite instincts.
Of course, it wasn't any kind of UI bug or automated moderation. The experience gave me a visceral understanding of what eventual consistency means - a term I also first learned around that time, during internship in an Erlang company, and connected the dots.
I was one of the first people in my social circle to spot the issue, but as it became apparent over a year or two before eventually getting fixed (or at least made less obvious), I ended up giving a very high level intro to distributed databases to quite a few non-tech people, in order to alleviate their concerns about Facebook gremlins.
This and other experiences using and building software systems make me agree with you. Especially for large web platforms, that weird thing you're experiencing could be some nasty form of shadow ban, but if you just noticed it after doing something, then chances are it's just a transient issue with queues or database consistency.
When they added the ability for Kerbals to be killed in Kerbal Space Program, they tripped over a bug where the first Kerbal in the game's engine, Jebediah (the one who dated back to the original introduction of astronauts at all, where only one existed), could not be killed. Because of some of the game logic having gone unmodified from the earlier versions, some operations would cause him to be loaded into the pilot seat and those operations didn't check if he was deceased. As a result, you could lose him on a mission only for him to spontaneously appear at the controls of another mission.
The community responded with fan-art of "Jebediah Kerman, thrillmaster."
When it strikes, it eats you swiftly, and eats you whole. It disassembles your crafts, spaghettifies your crews. No weapon or speech will save you. The Kraken transcends reality - even time travel, reloading from a saved game state, does not always stop the attack.
--
Of course, the Kraken is just a manifestation of unstable physics calculations and a floating point-based coordinate system: as you travel further out, the spacing between two consecutive coordinates gets bigger, until it overwhelms the physics engine and your craft disintegrates. And if you start doing crazy stunts, particularly involving high impulses or very fast rotations, the physics will break even if you're close to coordinate origin, due to rounding errors.
> Somewhere in this outer space, a gruesome death awaited, death and horror of a kind which Man had never encountered until he reached out for inter-stellar space itself. Apparently the light of the suns kept the Dragons away.
[0] https://www.gutenberg.org/files/29614/29614-h/29614-h.htm
The reason it happen is that the random number generator has only a small state (I think 16 bits) a limited number of possible rolls, and some seeds have a really short period, which limits it even more. Also, the seed is stored in your save file. It means that if you have a bad seed, you will always have a seriously broken RNG and the only way to change that is to create a new character. I think in later games, the same terrible RNG is used, but a new seed is picked each time you start the game, so you won't stay in the same "table". And while "charm tables" are the most obvious consequence, if probably has other, less noticeable effects on gameplay.
The weird part is that while it is obviously a bug as it negatively affects the experience of some random players, it is rarely referred to as such. It even persisted between versions. Some people even wrote tools to find out early on which table you are, and techniques to get the table you want, along with a variety of RNG manipulation exploits.
At one point I had 10 “vassals” sworn to me, and in our allegiance that came with the expectation that you assist and mentor those under you. Stakes were higher when our allegiance committed to being red dot PKs on a white server with a few other stronger groups in play.
The server emulation scene has come a long way since retail shut down, but even the most populous servers are a pale shade of this game in its hey day.
I look at this differently. I think the extreme cost of an MMO should force innovation. Take the constraints for what they are and run with them. Accepting that you can't do it the "traditional" way is the first step to figuring out a better way.
There is nothing that says a high-quality WoW killer absolutely must cost 10 figures to produce.
It was such a special thing before it was hit with a lot of road grading over the 15+ years of monthly patching that smoothed a lot of the rough edges that made it really unique.
(the walk isn't exactly random but it still works - the pathing heuristic is real-coordinate agnostic).
So it's a stochastic bug where any given occurrence can be explained as "you just have confirmation bias, you don't notice all the times Zygy Zebra gets attacked" or "you must have been closer than Zygy even though you thought you were a bit further away". Especially if the game has a first-player perspective which makes it harder to estimate whether you or another player is closer.
I think the best bug was the movement possibilities during spell casting from breaking animations. It created one of the most complex and amazing PK (PvP) dynamics of any MMO to exist.
The complexity of being able to move only so much to still get your cast off, and being able to slightly fast-cast or hold long delay-casts to "outplay" your opponents created so much depth to duels it was incredible.
Even after all these years I can still remember the Arc cast I would do, it was the keyboard combo: Hold Left -> Hold Z -> X -> tap/hold up to control the radius Then Hold Right -> Hold C -> X -> tap up to reverse the arc to return to where you initiated the cast so the spell could go off.
Man I miss this game, I hope someone creates a wonderfully buggy remake someday!
Bugs really made this game one of a kind, it's sad it would be so hard to replicate. The way he uses strafing and delay casting on corners to fight odds is amazing.
I still remember my hands shaking when I was in PK fights like this in my youth. The consequences of death made fighting so much more intense in this type of MMO.
So I created my own "algorithms" for things like randomly choosing which type of plants spawn in the world. It weighs them by the local frequency of each land-type. Eg if there's more swamp nearby, it's more likely to spawn cattails.
It's awful. Not intuitive. Weights have to be passed in ordered smallest to largest. Posting in hopes someone will correct my entire approach or point out an industry-standard way of doing these things.
Utils.weightedRandom(percents, types)
function weightedRandom(weight, outcomes){
var randNumber = Math.floor(Math.random()*100)
for(let i=0; i<weight.length; i++){
if(randNumber<weight[i]){
return outcomes[i]
}
else {
randNumber -= weight[i]
}
}
} random.choices(population=['A', 'B', 'C', 'D'],
weights=[3, 2, 5, 7])
https://docs.python.org/3/library/random.html#random.choicesPython has amazing math libraries. I use numjs[1] for everything, which is a port of numpy. I frequently wish I had use of Python's libraries.
function weightedRandom(weight, outcomes){
var total = sum( weight );
var roll = Math.random()*total; // value in the range [0,total)
var seen = 0;
for(let i=0; i<weight.length; i++) {
seen += weight[i];
if(roll<seen)
return outcomes[i];
}
}(This last pert is not an unsolvable problem, and in fact random algorithms should be unit tested. But it requires the right kind of unit testing).
https://forums.ddo.com/forums/showthread.php/264543-Wi-Flag-...
> Anyone who played Asheron's Call probably heard of the AI bug associated with monster Aggro where a person name "Wi" was forever cursed with getting the aggro no matter what.
What is curious is that it took them a long time to find. I’d think that -as long as you believe there is a bug- this should be fairly straightforward to spot.
People blamed this behavior for everything from loot drops, to combat outcomes, and to aggro mechanics.
And also remember that AC was before MMOs became massively popular, and outside of specific events and locations, there wasn't always more than a handful of players in a certain area for a given server where aggro mechanics like this would matter.
This is the power of continuous integration.
There's a certain amount of optimism needed to keep going as a software developer, and then there's the crippling amount of optimism that a lot of people have which makes for difficult team dynamics. CI says it doesn't matter if it works on your machine, it's red on a neutral box so fix your problems or it's not going into the release. It's much harder to ignore Jenkins than to ignore David.
People learn through trial and error that the Wally Filter works on bugs, so denial is their first and best defense. Prove to me there's a bug. I won't spend any time on it until you do.
But as a huge fan of the original AC, it's always fun to see it come up still.
I thought maybe it was going to be some weird DB glitch or something far upstream from the algorithm selecting which player to attack, but it was literally the logic of the very algorithm you would first look at if you were aware of such an issue.
This was a fun read though. Finding and fixing bugs like this is some of the most satisfying work we do imo, and no one outside of tech understands =)
Why couldn't they reproduce the problem in testing? For what was it, years?
Systematic flaws: a cross between groupthink, early flawed assumptions, deference to team leads, a 'I just look for 1hr, if I can't find move on' (which leads to not looking), or just plain simple "reading" instead of searching.
Lack of tooling: many game engines are infamous for lack of control over tooling. I havent used many, but I understand it would be quite an effort to run meaningful parameterised or structured fuzz testing on most systems. This makes it hard to artificially confirm suggestions. That said, there is practically no excuse for them not to just add a bunch of counters to the game - even on their internal testers it would very quickly become clear there was a bias.
Most of my 'should have caught it earlier bugs' are of the 'deference to lead' variety. I looked, didn't see immediately, handed off with some notes, and then the follow up debugger(s) took notes or thoughts as gospel. This is really hard to fight - I write something along the lines of "my hunch is there is a problem in code x because it handles y and is poorly structured/tested. I checked z and found i, j - queries as follows" and then find the debuggers effevtively refuse to look anywhere past x. This is particularly true for a group of debuggers, who play chinese whispers with groupthink and invent reasons it must be x.
Another hurdle is likely that game developer culture strongly favors integration testing over unit testing. Games are optimized for fun, not correctness, and you can't unit test fun. This specific roulette selection function would have been straightforward to unit test, and a unit test would have caught the distortion. But now imagine people keep varying how important distance is to the calculation in order to make it "feel right". Updating those unit tests is suddenly a noticeable slowdown on how quickly you can iterate on game feel.
But mostly it was to explain to people that a 1 in 100 chance doesn’t mean you’ll get it even after 200 goes.
Maybe the most occult part of this is figuring out that the unique IDs assigned to the players play a role.
Well, there is a larger bug -- the entire algorithm, functioning properly, still won't behave the way this letter says that it should. It's not clear what it's designed to do, but it's very obvious that it doesn't do what the description says it does.
Did nobody ever notice that?
> The problem comes up when we are assigning portions of the range to various players. If we wanted distance from the creature to be proportional to your chance to be selected--that is, if the closer you are the less chance you have of being attacked--then we would assign this range by taking your distance from the creature over the total distance--the distances of everybody under consideration added together. But we really want the inverse of this ratio
That couldn't be more explicit. In the example model, where distance is danger, player D is twice as far away as player A, and has twice the chance of being attacked.
To invert that, in the game, where proximity is danger, when player A is twice as close as player D, he should have twice the chance of being attacked.
The game's algorithm does not attempt to do this. In the worked example, player A is 50% more likely to be attacked than player D is.
The correct algorithm is not difficult to write or to execute:
1. Assign all players equal odds of being attacked.
2. Weight the odds by the ratio of (distance_to_furthest_targetable_player / distance_to_me).
To make the example easier to follow, assign player D [distance: 10] 60 units of probability space. Then player A [distance: 5] should receive 60*(10/5) = 120 units, player B [distance: 2] should receive 300 units, and player C [distance: 3] should receive 200 units. Generate a real number (or, heck, an integer) in the range [0, 680) and you have your selection. Or, if you prefer, normalize all the odds and then generate something in the range [0, 1). But how did they pick the crazy algorithm they're actually using?
I wonder if there’s a cheap Frankenstein approach to create a test-only interface into your code with a CLI / FFI and wrap it with some Python to easily test the result distributions.
Then you can use tools like scipy or newer probabilistic programming frameworks like pyro. Not sure what the FFI story is in R, maybe something similar could be done there.
Someone probably manually tested the feature originally and thought “yeah this feels random”
I had a character named Sal Monella, and I ran around the various major hubs trading people a single Fried Egg.
Good times.
Instead, the code assigned subintervals and then rolled a number between 0 and 1, instead of 0 and num_players. If your player happened to sort to the top of the list for subinterval assignment, you'd be it. You'd always be it.
Someone had to think to look at this, then confirm the maths, before it was found, years after the symptoms got reported. A unit test would have caught this, but didn't. Writing tests is annoying, especially in a codebase that keeps changing, but it's so important that this should count as one of those "this is what happens when you don't" lessons. Money was lost here.
It's slightly more complex than that: if you were farther away or otherwise in a better position than average your portion of the interval would be <1, and so even if you sorted first you still weren't guaranteed to be hit.
I've always wondered the best way to write tests for "This event should happen x% of the time."
Obviously we could re-run the test 100 times and see if it happened close to x%, but not only is that inefficient, how close is "close"? You'll get a bell curve (or similar), and most of the time you'll be close to x but sometimes you'll be legitimately far away from x and your test will fail.
You could start from a known seed, but then are you really testing the percentages, or just checking that the RNG gives you the same output for the same seed, which you already know?
1. Start with a known seed. 2. Run a single test run like what you say, and verify this run by hand. 3. Freeze the test in this state, that is, assert you get that exact result every time on the given seed.
What this creates is not what I would strictly speaking call a "unit test", but it does sort of pin the algorithm to your examined and verified output. In this case, a human would quite likely have caught this problem on a decent test set. Obviously, there are other pathologies that would slip right by a human; the human being careful only raises the bar for such pathologies, it doesn't completely eliminate them.
But at least freezing it solves the problem where a change you did not realize would be a change slips by unnoticed and this function suddenly has a completely different outcome.
This has worked for me, in the sense it has caught a couple of bugs that would have had non-trivial customer implications. But I've never worked on an MMORPG or anything else where randomness was intrinsic to my problem; it has always been incidental, like, is my password generation algorithm correct and does this sample of my data look like what I expect, not the core of my system.
I use https://github.com/pkhuong/csm/blob/master/csm.py to set a known false alarm rate (e.g., one in a billion) and test that some event happens with probability in a range [lo, hi], with lo < expected < hi (e.g., expected +/- 0.01). The statistical test will run more iteration as needed. If that's too slow, you can either widen the range (most effective), or increase the expected rate of false alarms (not as effective, because the number of iteration scales logarithmically wrt false alarms).
I can't remember the exact math I used at the time (I had to crack open a stats textbook), but ultimately it boiled down to generating a bunch of (seed, number of values previously pulled based on that seed, value) tuples, running a linear regression against them, and defining a maximum acceptable R^2 value based on my 10^-30 acceptable probability of a false fail.
When the RNG is not the thing being tested, mocking the RNG to do a sampled sweep through the RNG's output range is typically the correct move.
What would have helped in this case:
* Test that all players can be selected by the algorithm at all
* Test that the totals of the weights is the same as the RNG max range
Now a player named Wi is still the subject of conversation on at least two major online discussion forms over 20 years later.
From what I've read, it introduced quirkiness and a bunch of unintentional community involvement, with likely much hilarity.
The goal of playing a game is to have fun.
Dealing with frustration can also be fun, in fact the Dark Souls series recognized this brilliantly.
This bug, while undoubtedly frustrating, sounds like it was also a lot of fun for a lot of folks.
Unit tests for things with probability are hard. I've written one recently, and among other reasons, it convinced me to write the selection differently so it was both more 'fair' and more easily tested.
To generalize, it's vital to have something like Sentry from day one - a low-cost abstraction that lets you monitor broken assumptions asynchronously. Though of course, these kinds of tools didn't exist in 2002!
The first unit testing library, JUnit (Java), was released in 1997.
Asheron's Call was released in 1999 after four years of development. It's quite possible that this bug was introduced before the concept of unit testing even existed or was widely known.
It's also possible to write a test for this where the behavior in the test accurately tests the wrong behavior. The error is pretty subtle. There was even a “bug” in your summary of the bug ;)
Only if you don't have a functional core.
One of the psychology tricks I've learned about unit testing is just how bad sunk cost fallacy affects some people (all the time, and most people some of the time). I've witnessed people pairing for two days to fix a couple of nasty integration tests. That's 3 person-days of work lost on a couple of tests. How many release cycles would it take for that test to pay for itself versus manual testing?
You want the tests at the bottom of your pyramid to be so simple that people think of them as disposable. They should not feel guilty deleting one test and writing a replacement if the requirements change. Elaborate mocks or poorly organized suites can take that away. Which is hard to explain to people who don't see the problem with their tests.
You also want the sides of that pyramid to be pretty shallow, especially if you keep changing your design.