The Outer Worlds: Fixing the “game thinks my companion is dead” bug
twitter.com
twitter.com
> that was our fallback, but
> -without actually knowing the cause, we couldn’t be sure that’d fix it
> -whatever the cause of the problem is might be causing OTHER problems, so we should find it
> -games near ship are fragile and it’s best to not make changes you don’t 100% understand
_THIS_ is someone that is practicing good software practices. I have no doubt this is why Outer Worlds is so much less buggy than, say, recent Bethesda _patches_. I saw a report this morning that the latest Fallout 76 patch causes some armors to degrade when a weapon reloads. This sounds exactly like the kind of bug you'd get from "fixing" something without understanding the actual cause, and thus just kicking up more emergent bugs because of your fix.
Complex interactions (like modern games), filled with shortcuts and assumptions (like modern games) means some bugs are inevitable. Requiring rigorous understanding before making changes can keep those bugs from multiplying.
...says someone with zero experience in coding games, but plenty of experience both fixing and generating bugs.
[Edit: fix formatting]
This is the antithesis of defensive programming. It's a game-breaking bug that's allowed to happen because of two things:
1. They allowed companions to take damage because they marked them with the wrong property
2. They never checked if the companions went somewhere silly to warp them back to where they should be
That's not defensive, that's like calling an API without a defensive try/catch if something goes unexpectedly wrong.
If they'd have done either of these things (and sent back error reports when it happened), they'd have caught it way sooner.
I don't write games and appreciate the interactions can be very complex. While I can understand your delight in them finding the root cause, I can't understand your support for failing to put in some defensive measures when they couldn't recreate a frequently reported bug.
Defensive programming is not fixing bugs with more features, quite the opposite: it is failing hard when a condition does not hold.
If it shouldn't throw, don't try to catch... That is an unexpected failure that needs to be understood.
Thank you for standing up for confident programming! (We could make this a new term if we wanted to :D)
EDIT: Wait, maybe there's a better term than confident programming already... I would appreciate some input if someone sees this.
The best solution is, as always, context and resource dependant.
At the same time, you really should try to validate inputs. The idea being that most exceptions caught should be of the form 'Input was invalid'. Rather than 'function crashed'.
> without an understanding of IF it should and WHY it might
rocqua explains it best - the API handler is an "input" to the overall system. That input may be "invalid" in that it fails to handle its own errors. In that case, you should catch a general exception rather than crashing. In this case, you know that it might raise an error, why it might raise an error, and that it's OK to ignore (well, log/alert/whatever.)
The situation I was responding to is very different - it's basically suggesting rather understand what the API you're invoking may be trying to tell you, just try/catch anything it throws and ignore it if you didn't expect it.
On the other hand, a system where damage is prevented in some ways and assumed not to happen in others is already an overly complex construction of interacting systems. Switching to an invulnerability flag could simplify the system.
In a modern open-worldish game, this is pretty nontrivial.
This is exactly the sort of bug that anybody under the kind of bone-crunching compression you see in game development is gonna fix after they see it. "Don't start bandaging until you bleed."
Game engines do everything every frame by default, and the game has to disable/enable stuff from happening, which leads to bugs like this. With a more functional approach instead of tweaking the engine state these bugs can be avoided.
The game's hardest difficulty setting - "Supernova" - allows companions to die [1] (as well as a bunch of other changes that make the game tougher, not just stronger enemies)
[1] https://www.pcgamer.com/uk/which-the-outer-worlds-difficulty...
From the tweet series: "There were one or two cases before launch where this issue seemed to happen, but no one in QA ever managed to reproduce it and despite our best efforts we couldn't learn anything concrete about it"
Assuming this is true, before launch, it was NOT a frequently reported bug.
After launch, they dove in and figured out what was happening AND what triggered it.
Your complaint appears to be that they (1) had a bug, and (2) didn't resolve boundary checking in the production code. Given the number of "fall through the map" bugs in countless games, I assume #2 is, in fact, hard. Writing non-trivial code without bugs (to avoid #1) is also hard, based on...all the smart coders I've read.
Regardless of the source of the bugs, I'm cheering that they didn't just put in defensive measures and called it a day. To the point that defensive measures can be a source of bugs, consider how many "violent vibration" bugs we've seen in, say, Skyrim, Fallout, and GTA. (While I don't play, based on clips I've seen from GTA and Red Dead, those engine(s) spam duplicate entities when the boundary overlap is detected, where Skyrim/Fallout just vibrate and rattle). I find it very plausible that attempts to "warp" characters having positional bugs without understanding the source of those bugs will end up a source of new bugs.
Can’t gamedevs reinvent solidness and make underground/walls solid? Viscosity/density/archimedes?
RAGE/GTA/RDR2 are so insanely buggy games that a game developer probably can learn nothing from them, except how to not do it, and that no matter how broken your games are, you can still earn billions.
Now you created a whole zoo of new troubles, including a whole set of troubles you cant even predict, because you didnt know about them when you made this giant default behaviour case.
if (specific case){
} else (everything else that could ever happen, inlcuding all future edgcases){
// here be dragons
}
Examples? Leveldesigner adds elevator, that transports player and npc ocassionally outside of the teleport box.
Skybox entitys are npcs that live outside of the box- so space battle npcs, get telported to bunks..
etc. etc.
Handling the unknown - with any other paradigm then the fail early, fail loud - is the root of evil.
Take all the time you think you can afford if yyou software is not released. But it is being used by customers and such a large bug is in it and if you are not reacting to it then imo it is a problem.
I think people disagreeing with me is unable to see from the perspective of customers. In this case a player that loses hours of progress because of a bug deserves some atention
> If there was an "easy" fix that can be applied in a day or two,
> it should have been done. Agreed.
However, when someone asks for this, the implicit assumptions are that the "easy fix"
- actually fixes the symptoms (even though ignoring the root-cause),
- and does NOT introduce additional unexpected behaviour.
Ensuring this may sometimes make the easy fix just as time-consuming as a proper fix, thus favouring doing it once properly.
There's nothing specific about software in there. My mechanic does the same type of reasoning.
Though in defense, most of them are good enough to guess right 90% of the time. And mechanics generally get paid by the task, not hour. Replacing the probably misbehaving X pays. Spending an extra half an hour making sure that X really is broken and replacement is the best solution doesn't pay. So they just replace the X and hope for the best, and then replace something else if it didn't work.
That we do have to worry about breaking something in the software industry field, just tells us how far we've yet to go.
In regular programming scopes are small, small enough for programmers to accomodate all expected logic. In games interaction scope is unlimited, or atleast that is what the game designers want.
If you try to program every expected system interaction you would never finish. So instead game programming is about making systems flexible. For example, as far as the gameplay code is concerned the inside of players ship is just yet another level. The bounds of any level are defined by what the map designers have layed out, not any pre-planned "ship area".
There are no meetings where project managers request "more playable area for level 7".
Likewise the systems programmer who made the furniture system was not the script writer who created the ladders. The script writer likely thought it was brilliant and elegant to split the entrances and exits, because in game programming it was indeed brilliant. It allowed the save system's existing code to handling what would have otherwise been special handling. It allowed the existing combat logic to handling what would otherwise be special handling.
Of course regular programming is different. In regular applications if a user's data is being sent to a server there is no secondary system which needs to be allowed to kill this data. Nor is there any save system which needs to be able to save and restore an in-progress action.
Games programming is all about figuring out ways to allow more varied and more complex interactions without exploding the implementation exponentially.
Yes regular programming could indeed add a system where numbers might be animated. During the animation maybe you want another system to be able to change the number. Maybe the number disappears, make sure you're front end can handle not having the requested data arriving. Maybe you want systems which add extra fields to user data, maybe you want anytime a user attempts login that there is a 1% chance they receive a reward determined by their "user rank".
All these are "possible" but do not happen in regular programming. If a user attempted to log into your application, and instead a window popped up saying "you have died from starvation, please recreate your account and play again" rightly so your users would be angry, not impressed.
Different constraints and goals lead to different priorities. Games on a per screen basis have a lot more budget available. Imagine if your application's login screen took 2 months to program and required the full time effort of multiple artists.
IMO it is more useful to think of software and its development as simply more or less complex.
Also, your whole post basically ignores GP’s argument that not everything is either a CRUD app or a game. Maybe it would serve your argument better if you took care to read and respond to his.
It’s hard especially if functionality gets added mid process (the companion can die now).
I had a piece of code log (“you should never get here”) once before crashing. Some programmer long gone put a constraint check with that strange note.
It’s very hard especially if functionality gets added mid process (the companion can die now).
I had a piece of code log (“you should never get here”) once before crashing. Some programmer long gone put a constraint check with that strange note.
This is not the same kind of thing as bank software.
Originally, when I read the article, I thought that the issues described may be a result of bad software design.
But maybe it’s because applications are about strictness, and games are about freedom.
In application programming, you should be in a defined state most of the time, and make state transitions short and atomic.
In guess that in games, you are really in one huge, never-ending state transition.
The type of programs where things like this occur are called simulations.
Much of what player consider progress in games is the industries ability to add more complexity. If you look at the contemporary equivalent to chess that would be the spin-off of MOBAs: AutoChess.
AutoChess may bare a similar name to chess but it is a vastly more complex game. In theory AutoChess could have been created on the NES, but no one would have understood how to keep the implementation reasonable.
Instead it took a progress, from Chess we got turn based strategy games. From TBS we got real time strategy games. From the early RTS games, we got the story rich Warcraft. From Warcraft came extensive modding support and story based characters. With said mod tooling and characters a series of mods became what we now know as LoL and MOBAs. Another decade and another set of mod support brought AutoChess.
Yes simulations are complex games, but what used to be a defined genre is now an aspect of most big games. Western RPGs certainly try to push the boundary of coplexity, as to tycoon or simulation games. Still, all genres are tending towards more complex systems.
In that particular case, what would be your favorite fix. I tend to say that the second step of climbing a ladder should probably not count as a new interaction.
For example, if you're standing in a doorway talking to someone and an NPC wants to close the door, it could close in between you and the NPC you're talking to.
It would be hard to make believable interactions when unexpected environmental side effects can happen at any time
They should probably also consider adding fall damage as part of the considerations in invulnerable mode.
Edit (additionally):
If they can't be made atomic, make the transitions timed, with expiration, never allow deadlocks to occur and always branch to a fail-safe state. Players will tolerate things like the uncontrollable NPC getting pinched off screen and respawning, even with a comical animation if it's that kind of game; they hate a blind referee ruining things for no perceivable reason.
In the case of another game where someone's actually making progress up a very large structure, I can't think of any game offhand where saving in such a state has actually been allowed. However allowing a save there by necessity triggers the edited version of my statement which includes a timed state transition and safe aborts for failures in state change. Such a save would also include the current state of progression as well as the remaining limit for the timer.
Probably the most ludicrous is that when sneaking, only the player is taken into consideration by enemies. I've seen patrolling enemies kick my companions out of the way without actually noticing them.
Reading the player's intent would probably still be a very difficult problem, so even with the perfect knowledge of how to interact in the world they'd probably not act and react to changes in the player's intent anywhere near as well.
Skyrim was/is awful about that. I'd be sneaking along through a tomb, and then Lydia screams "I am sworn to carry your burdens!" and charges in, waking up a dozen draughir.
When large amounts of programming time is spent on game ai it does end up "self driving" at an adequate level, see Google's sc2 bot, Google's older arcade game bot's, and openai's dota bot. It's just not worth it investing that amount of time into most games.
They generally move at a pretty fast pace and change often(due to the iterative nature of what's "fun"). I don't think I saw a single unit test until I was well out of the game industry.
Their blog regularly goes into details about how it all works. It's very much worth reading.
It's been a few years but if I recall correctly we did have a "smoke test" running on a jenkins box under someone's desk that booted a few levels with most of the game entities(they were called danger_room_1/2/3/etc) and verified that the game didn't crash. Aside from that though there wasn't much else which lead to some spectacular build breaks.
Ah the joys of "AAA" development.
Not letting companions fall to their death, ever... Bethesda hasn't fixed these bugs in 5 years.
What would you do instead?
These problems arise when a game engine starts as a generic physics simulation and various restrictions/rules mandated by gameplay design are added as an afterthought.
That might make sense when you view it through this specific lens, but that design decision allowed them to save character positions on ladders and make 'getting on' and 'getting off' be separate 'things' - just off the top of my head.
The '2 interactions' method actually makes the most sense in the context of their game engine - it's the same way getting 'on' or 'off' a chair works.
Their fix of just allowing interactions is probably the most correct option. That is essentially toggling a setting, as opposed to re-engineering how ladder mechanics work.
Games are complicated.
Altogether an elegant solution for managing what is otherwise a special case.
If companions should be unkillable, then don't count on eliminating all sources of potential damage. Really make them unkillable. A ladder shouldn't ever take someone higher than the ladder goes so if some guy attempts to climb higher than the ladder goes, then freeze them in place or make them climb in place. Since these shouldn't trigger, crash the game if in dev mode, and if in prod mode check whether player allows telemetrics and if so phone home.
Those are the failsafe fixes.
Finally fix the root cause here by special casing that getting off a ladder during interaction is allowed.
Reproducing it was hard but the culprit was WiFi interference. The protocol had separate on/off commands and no provision for lost packets since it wasn't natively designed for wireless. If the off command dropped out everything would act as if a button was still pressed.
>_<
> I had the same problem in a hardware product using a legacy wired protocol that was sent over a ZigBee connection.
Just speculating though....
One reason it was so hard to pin down is that it was impossible to tell when the bug actually happened -- all of the cases we had were essentially "hey something bad happened in the last ten hours and now my quest is broken" (5/18)
Depending on how early they realized it was limited to companions on the ship, with a specific state (the source is not clear on how early they determined that to be) they could log events that damage a player when in that specific state.
That would have gotten them to fall damage, at which point they might be able to make the leap to ladders. If not, they could then focus their logging around changes in elevation when in that state.
But that's backward reasoning from the already known end state, so it's not fair to apply in hindsight. It's also not clear how much log data they can ingest, whether end-user telemetry was feasible, etc...
Turns out the flaw in the logic was that the NPCs can reach great height on the ship. But the flaw in the logic just as easily could have been in the combat. (Or some other way that the tweeter didn't mention).
In any case, it's really fun to read the "aha! got it" moment.
> The only place in the game when a companion is present but not in the active party is when the player is on their ship
> Eventually we figured out that "undamageable" does not mean "invulnerable" -- they can't take damage from attacks but can still get hurt from other things
> One of those things: falling a great distance
> The problem with that is that there are no spots in the player's ship that are high enough to result in a lethal fall
> So now we had to figure out how companions were mysteriously ending up way above the level
> I looked into tons of theories
This does look like logging would have caught it. They knew exactly what they were looking for: companions reaching strange heights while the player is on their ship.
[1] https://llvm.org/docs/XRay.html#flight-data-recorder-mode
But it's difficult to police good telemetry and easy to change silently. Bad actors have continually given justifiable reasons for data collection, sometimes even with a pinky-swear that they'll not use it for evil. But the terms can be updated by the party collecting the data at any time they desire, with little recourse for the user other than to stop using that product or service.
So no, I'm not worked up over that kind of telemetry, but I'm still going to block it. It doesn't make sense not to.
If it were opt-in, it could be used well however, similar to what the Telltale games had done.
65% of people made choice A. 35% of people made choice B.
They have a sort of proxy for that with Achievements however.
Many bug reporting systems already work this way, such as the one used by Firefox, I think.
Just because something can be done by violating user privacy, that doesn't mean it's the only way, or even that alternatives are difficult.
The problem is that devs reeeeally need a mix of very specific data for debugging, and very aggregated data for prioritization... There's all kinds of cool differential privacy tools for doing things like this, but no way for users to know or trust that they do what's claimed on the tin.
I've been floating the idea of setting up some log shipper plugins for Unity/Unreal to make a turnkey solution you could drop in and self host.
I will say that TOW was very stable for an Obsidian Game. I don't think I've experienced a bug yet.
The Outer Worlds is an unexpected joy, too. Definitely an open-world game I had no idea I needed. If you haven’t, check it out. It’s basically Fallout 4 but in space.
Notably, I found that the Vicar had a really interesting progression as a character and his relationship with religion. It lead to some interesting philosophical dialogue, especially if you bring him when you meet the Iconoclasts. He has a healthy debate with their leader and they both come to the conclusion "let's agree to disagree". It was great.
"The game, initially developed as a joke prototype from an internal game jam and shown in an early alpha state in YouTube videos, was met with excitement and attention, prompting the studio to build the game into a releasable state while still retaining various non-breaking bugs and glitches to maintain the game's entertainment value."
You can launch yourself into the air (100s of feet), run inside a dumpster to move it, drop through the floor to start falling from the sky, booby-trap a trashcan so it flies out of no-where and hit you upside the head, walk up an invisible ramp... it's fun.
In addition, there's the Super Mario 64 0x A Button challenge, an effort to beat the game while avoiding the use of its main mechanic. It's not a speedrun, but the themes of exploiting every glitch possible is the same. It's also the source of the infamous "0.5x A Presses" video. https://www.youtube.com/watch?v=kpk2tdsPh0A
Also this sub that highlights small, easy to overlook game details: https://www.reddit.com/r/GamingDetails/
https://kotaku.com/guy-beats-fallout-4-without-killing-anyon...
> The only logical culprit was a bit of scripting that runs when a companion's health reaches zero: if they're in the party, it waits for combat to end and revives them; otherwise it marks them as dead "for real
The code should never have been written this way if the companion is not supposed to die. There should have been an `is_dead()` function that always returned false.
This site has declined dramatically and I think it's a result of the quality of programmers we have today vs what we had 10 years ago before the internet became mainstream.
None of these comments nor the tweets talk about how bad that code was, but instead kiss the guy's ass.
I expect to see more and more posts about games and other trivialities and less and less posts about important topics as the years fly by.
First of all the fact that the game could be finished in 12 minutes was a surprise to both devs.
And at several points during the video dev#1 expressed surprise over a shortcut the player was taking, while dev#2 said he did that in testing all the time.
Seems like a lack of both QA and communication.
Projects are often rushed to completion for shipping, and open world games can't be the easiest thing to test. You're balancing player freedom and bug hunting. It's a challenge I can only imagine.
I don't think the fact that developers are surprised by speedrun times means that a game is necessarily bad! (And even many of the most extraordinarily beloved games have glitches, often because realistic physics simulation is hard.)
I think I saw a video of Bennett Foddy in which he noted that other people could complete "Getting Over It with Bennett Foddy" considerably faster than he could, and maybe faster than he imagined anyone would be able to.
In good software, the bugs don't cause user-facing issues. QA teams aim to make software good by finding them.
Speedrunners aren't users or QA. Speedrunners find bugs no matter what. From the developer's perspective it really doesn't matter if you can save 20 hours beating a game by clipping through the wall next to the boss, if that clip is so immensely complicated to perform that no user would ever encounter it naturally. The devs can be surprised, or not surprised, to hear their game was beaten much faster than they anticipated by speedrunners; either way it does not make their game worse or rushed or a product of bad communication.
The Legend of Zelda: The Wind Waker came out in 2002 and sold 4.6 million copies. In July of this year, a new bug was discovered by dedicated minds, not randomly but after nearly two decades of hard work, that saved two hours on the speedrun by clipping through a barrier that was supposed to be unpassable. Does it really matter from a quality standpoint if this bug exists or not? Would your original rating of The Legend of Zelda: The Wind Waker be lessened 17 years later because this bug was found by people who perhaps spent more hours trying to find it than you've spent playing video games in your entire life?
A: That is true for all games
or
B: You're holding everybody to too high a standard.
Some recent examples: There's a glitch in Zelda: Breath of the Wild where if you shield surf onto an enemy while it's taking damage while you're in slow-mo because of shooting an arrow while falling, you'll bounce off of them about a mile into the sky. In DOOM (2016), there's a physics bug that lets you rocket into the sky by standing on a railing and then... standing up. You can also glitch through doors by triggering a specific take-down animation on zombies.
Those are incredible games, and relatively bug-free.
The game can be a 30 hour game if you hussle or as far as 200 if you don't.