Tough times on the road to Starcraft (2012)
codeofhonor.com
codeofhonor.com
I'm having a version of this problem right now. One of our products is not going to be sold to new customers at some point soon (but will continue to have plenty of existing customers), so it receives a small-ish budget for maintenance.
A lot of people seem to believe that this means any problems should be patched over with duct tape at the cheapest possible up-front expense, even if it comes at the cost of increased complexity.
I think of it the other way around: we are going to be forced to maintain whatever complexity we add to this with a very small budget, so at this point we shouldn't make any change that increases complexity. Ideally, we'd only make changes that decrease complexity! Even at higher up-front cost.
But it's hard to get people to understand the logic of that. I might not be phrasing it well.
It's a little like business people and developers think that most of the budget of an actively developed product goes to new development and only a small fraction is needed for maintenance, when in practise it seems to me to be the other way around.
In my experience, I'd go 10/90.
- Sometimes we are fully convinced that the new feature/product is going to be used and generate revenue and that we will have to maintain it for months after ship date. In this case it's often worth it to do things "right" as a bit more time early can save a ton of time afterwards. This is typically the case for incremental updates of an existing product with an established client base.
- Sometimes we have no idea on the success of the product/feature. Doing things "right" is far less critical that just shipping, as their is a non negligible chance that you are just unknowingly wasting time on a useless component. Doing a useless thing "right" has no value, spending more time on it is pure waste.
You often see "horror stories" written by engineers about products that were rushed too fast to prod, became a hit, and had to be significantly re-engineered. Business wise, these actually are success stories as the business was able to create a successful product, sell it, then improve it. Even if things were a mess behind the scene.
You also sometimes see engineer being proud of winning a fight with management in order to improve a product before shipping, and then having that product being a massive success. Take care that there is significant survivorship bias here: you almost never see the same engineer writing about the product he perfectly engineered, fought to push the release date and that got 0 user. He'll consider that the lack of business success has nothing to do with him. These things happen all the time, you should consider that the default state of any new significantly innovative product or idea is to generate a total of 0 revenue, and you should act accordingly by helping the business validate the idea as early and fast as possible. And sometimes that means taking on technical debt.
A great engineer will be able to accept that estimating the success of a brand new product/idea/feature is extremely hard and adapt his approach depending on the situation, accepting debt in the process if needed, going to prod "too quick if needed etc. The same engineer can also afford to be ruthless in asking for time to do things right when the business aspects have been validated because he has shown his ability to compromise and understand the context in which he operates.
That said, I'm not married to the idea of getting it working at all. I'm perfectly fine with scrapping it entirely. That's a business decision I currently don't have the data to judge on.
All I'm saying is that this is something where if we are going to build it, we know are going to have to support it for a long time into the future. So if we fix it, it should be fixed at low added complexity. But it's also an alternative to declare it dead, of course.
1. Doing things right often results in releasing faster. But wildly unrealistic schedules sometimes cause teams to reject the right thing infavor of the "quick" thing you pay for ten times over before release.
2. The word "debt" implies that it can be paid off. In many cases, removing the complexity introduced is simply inteactable.
3. Complexity has a huge cost throughout the entire stack -- testing, building, fixing bugs, etc. A shortcut may just be a shortcut, but if it introduces additional complexity, that's qualitatively different and you need to be careful.
Basically spent ~3 years as the sole developer on an application in "maintenance mode" that kept getting more customers since it supported a bunch of features the intended replacement didn't. That increased client base wanted new features in the maintenance product that leadership wouldn't turn down. The product still maintained a turn-around time an order of magnitude less than the primary replacements.
A few days before I tendered my resignation, the company had laid off about half the development team for the primary replacement (even though they had contracts promising work that was planned to take 12+ months with the full dev team), and declared a different product as the "primary replacement" for the other ones they held. Don't know how that will work out for them, don't really care either.
New place has a standing policy that basically says if we can prove a maintenance system is taking more than a few hours a month, they'll authorized repairing/replacing/removing said system. It's incredible how much of a difference it makes on employee morale not needing to constantly context shift to spend a few hours fixing some broken system over and over again.
Why does this always seem to be the case? If someone sets a world record in sprinting, we don't expect them to exceed it every year. Yet in business, if you pull a miracle in the 11th hour at your job, that becomes your new baseline next time reviews come around.
Now Warcraft III, which came out only three years later, had a fully deterministic game engine and, for replay files, only saved players inputs. Hence the Warcraft III save file, even for entire games, were tiny. And there were several websites where you could download and replay save files from other people, including from famous matchups. Fun times: I'd exchange my best games with my brother, as email attachments IIRC. We'd then watch each other's games and make comments.
IIRC Microsoft's first Age of Empire, which came out in 1997, already used a deterministic game engine so there was already an AAA title who used that technique when StarCraft came out in 1998.
That's interesting. I wonder why the other poster say the replay functionality would turn into nonsense at some point.
Now if the game is deterministic and multiplayer works by only sending user inputs around, why the save / continue feature wasn't implemented that way?
In any case I love to read about how these games worked!
Your only interpreter for the data is the entire game engine. By design it's going to use most of the power of the target hardware, so cannot run faster. Ie, 30 minute load time if you had played 30 min.
- a huge part is going to go to rendering, which you don't need to do if you only care about the end state
- in many games, the simulation runs at a much lower frame rate than the game, so perhaps you see the game animated at a silky smooth 60 fps but internally the sim runs at a fraction. 10 and 5 fps for the sim are common.
- some sim architectures may allow you to run the sim at a variable frame rate and still generate deterministic results. So (ideally) you tell the sim to run a single step of 30 mins and you get the end result. In practice of course variable steps will be limited to much less, but Supercell mentioned reducing their server sim cpu costs by 80-90% with this kind of trick.
All added up with reasonable guesses for a well optimized case: 30 mins of sim run at 1 sim tick per second of game time, where sim ticks take say 2ms, could take 3.6 seconds to sim, not 30 minutes.
The problem is floating point calculations on different processors. Maybe in the 90's that wasn't such an issue, but good luck now. Plus, if you don't need a physics engine, maybe you can stay in integer space.
That would explain small deviations at first, and getting bigger by time.
For example a game like Braid also chose not to use a deterministic engine for that reason.
Haha, I'd be surprised if they used floats at all in a 90's game. Fixed point were popular.
Edit: this is it https://www.youtube.com/watch?v=HrUF-LFTs-A
The whole talk is very relevant to anyone interested in this topic.
A 3D engine in fixed point (which I never fully finished) gave me a lot of headaches. So glad all those devices have proper floating point and a gpu's.
> Originally we were quite afraid of Floating point operations discrepancies across different computers. But surprisingly, this hasn't been a big problem so far (knocking on the table). We got away with implementing our own trigonometric functions.
I usually don’t bother since I don’t care about the whole replay and don’t want to wait. But it’s handy when you want to learn the strategy of somebody better than you that you just played against.
Civilization 7.9.9D was my favorite custom back in the day :)
There is a major drawback to save files which rely on the deterministic nature of the game: what happens when the determinism changes? I don't mean due to bugs; I mean intentional modifications to the game's behaviour through patches. Either these changes will break old replays and saves, or the game needs to carry around multiple simulation versions.
You bring up Age of Empires. This is a small problem in the current AoE2 competitive scene. Yes, AoE2 has a competitive scene – the game has been re-released twice in the past decade, and still enjoys the attention of competitive play, including sponsored tournaments (e.g., the Red Bull Wololo series). Microsoft are attempting to give the game the same treatment as other modern, competitive titles, which means frequent balance patches, and occasional injections of new content (DLC). Unfortunately, each one of these updates breaks old save games and replays, which, as you point out, rely on the deterministic nature of the engine. It means that any casting/commentating of games must happen within relatively short order of them being played, and that the only reliable distribution format for old games is... screen recordings. I'd much rather a less space-efficient save format (or at least having the option to convert deterministic replays into a fully self-sufficient archive format) than having to discard native recordings of great games by the wayside in favour of videos.
StarCraft actually let you save replays which were just player inputs - which begs the question why save games couldn’t do this, but I digress - and I clearly remember attempting to watch a replay after a major game update and seeing it fall apart after a few minutes.
The replay was of a 40 minute game. It ran fine for the first few minutes, then I’m not sure what went wrong but suddenly all the characters stopped moving (except the automated mining drones) and the game just stayed like that for the next 30+ minutes.
Replays weren't added until 1.08[1] so it may have changed sometime post launch, but I'd have thought because it means you have to run the sim to get to the eg. 30 minute mark where the save is, and even on modern machines running remastered that takes quite a while. A more constrained machine from the times wouldn't be able to do the x16 speed fast forward used today.
In other words, as usual it's all about tradeoffs.
The ideal system would be similar to what we have with video algorithms. Every so often we draw out the full data needed for the existing state, then from then on-wards is just diffs until the next key frame is reached.
Also, it's fairly tricky to make the deterministic replay works. SC definitely do that for network play, but the risk is arguably higher for saving. StarCraft there were lots of synchronization issues even at very last minutes of its beta due to unexpected non-determinism which is probably because the game engine was not designed with that in mind. IIRC, the game would simply shutdown the session in a such case. But if it's from the save file, you don't have a way to detect that so the save file is going to be in a corrupted state with no hope of recovering. In the case of WC3, the engine has been written with lessons learned in the hard way so it should be capable of handling this problem more elegantly.
Every game programming story like this makes me feel like game programmers are the most undervalued programmers out there. Loads of them have very little experience, are asked to build extremely complex systems on fast moving ground, and get paid peanuts to do it.
I cannot imagine a system harder to put together than a game (maybe like... accounting software with a bunch of extra business rules?). Just so much stuff that can go wrong.
Can confirm. This is pretty nasty too. It's really tradeoffs all around.
Personally, I enjoy working on custom/DIY game engine code to unwind from all the banking crap during the week. I have come to terms with the reality that my side projects will probably never be finished in any meaningful way, but at least I feel like I have total autonomy over some complex thing.
Eventually, I may move back into game dev full time. For me it was never about the money. It is about doing something no one else is doing and having fun along the way.
I think financial independence (not necessarily wealth) is a big part of not letting game dev consume your soul.
I can't wait for the promised followup
Anyway, even if SC2 has its issues and by many regards can be considered inferior to broodwar from an esport perspective, saying “sc2 is just death ball” only reveals the unculture of the speaker.
At the top level players should keep an eye on the other player at all times, because a lot of strategies has a very specific response or you die.
With those MMR, you'd... probably not expect it to be beat Serral (7.5k MMR lmao...)
SC2 generally does not encourage significantly splitting up your main army in mid-late game. So ya, main armies are still normally moving around as a single blob/body.
I think every race still has enough tools to punish a-move, stacked/bunched up ground armies (lurker, baneling, widow mines, tank/liberator, psistorm, disruptor) that even if a ground army moves as a clump, there's significant pressure to quickly micro and split on contact.
There was a recent-ish period where there was a lot of dissatisfaction about late-game protoss air basically being a death ball, especially against zerg. I think that has somewhat dissipated - they released a patch that made disincentivized getting to that state, but I suspect that dealing with late-game protoss air is still annoying as heck for zerg.
It's easy to think that you were (in a sense) hired to type out that code over and over to realise the vision of the designer. If nobody tells you otherwise, you'll be more and more convinced of it for every time you type the code out.
> Among its many features, storm contained an excellent implementation of doubly-linked lists using templates in C++.
In such a context, the only cost of not creating a function is that you make the code harder to refactor... hardly a concern for a game developer.
I wonder if devs at the time were skeptical of relying on automatic optimizations, when a few years earlier doing it themselves was the only option. I could definitely see myself falling into that mindset.
Game dev never changes
There, I corrected it for You :)
My first solution involved something I called "agents" which are really just delegates. Agents could be chained, allowing a given behavior to be written just once and combined with other behaviors. These days I am fully on board the ECS train in my game work. Though it presents complexities of its own, ECS really helps manage the combinatorial explosion of state and behavior a game object may have.
https://www.codeofhonor.com/blog/the-starcraft-path-finding-...
Now: "we need to outspend on marketing and advertising and out-execute on in-game monetization."
Clearly, engineering is important. But games that are too well-engineered tend to be not fun. A lot of what makes a game fun are the exceptions to fundamental assumptions/rules/constraints.
I still love the game and would like to let my kids play it one day.
Every time I hear these stories from game developers, I just shake my head and think “where the hell is the project manager?” It’s as if the whole game industry just doesn’t believe in rigorous, disciplined project management, and instead always relies on crunch time, 48 hour days and an endless supply of naive youth to burn out. Nobody benefits when you keep saying “we’re two months from shipping” for 14 months.
https://en.m.wikipedia.org/wiki/Release_early,_release_often
So I would expect otherwise. Maybe 1998 was the tipping point?
They would do stuff like downplaying all the problems and ship a broken game, or pulling developers into redundant meetings to waste their time.
A software engineer that is proficient at memory management, multithreading, pathfinding, networking, graphics, linear algebra, computer science, data structures and algorithms, and all the other stuff involved in programming a game, doesn't want to listen to a person that not only doesn't understand any of that, but feels entitled to have an opinion on how tasks involving that should be triaged and executed, and sometimes are insufferable people obsessed with hierarchy and domination to compensate for not understanding things.
Want your project to be fast? Get out of the way. Let engineers collaborate and figure it out. And most importantly, stop trying to see software projects through the perspective of oversimplification.
Good project managers exist, but most are not. Hear it from Steve Jobs if you need to: https://www.youtube.com/watch?v=fj0hpsJvrko
Project managers set priorities, communicate requirements to developers, and communicate expected timlines, problems, needed changes, and other developer-surfaced issues back up the chain.
If engineers (like me) are left alone to work on it, either one of them becomes de-facto project manager, or you probably end up with a well-engineered product and wholly unsatisfied customers.
Stubbornness, cognitive dissonance are my explanations. I witnessed quite a few projects, whose deadlines got postponed for more than 6 month. Now comes the exciting part. These managers had a track record of postponing or more neutral: experiencing delay. Instead of fewer crunch time during the next projects, they stick to their previous behavior instead of trying to learn.
There you go, learned self-inflicting pain.
Investors and customers stick around because the company promises that it's just around the corner. Then that's unrealistic and the delays start racking up. (This is at least what I have seen, when deadlines come from all the way from top management and everyone below tries to squeeze their plan into those deadlines.)
Somebody benefits alright, just not the developers who are on the death march.
I did call it my self-aggrandizing blog at one point. In retrospect I think I mostly wrote it to vent about the pain and awfulness of crunch culture. I'll seek a therapist in future, and hope you'll forgive my trespasses.
Speaking about your own achievments does not lessen my achievements.
Thak you for the post I found it enjoyable.
Warcraft/Diablo/Starcraft blew my mind as a kid and were a huge part of my childhood, they played a major role in getting me into computers.
I don’t mind if you’re a little boastful.
Thanks for all the fun!
I was actually excited to see your blog domain pop up on HN today, only to find out it was to an old post. You have a ton of valuable insight and knowledge into games and project management and I hope that commenters like the one you replied to don't put you off from sharing that with the world.
I actually really wish we had our hands on some of the earliest builds, even the "orcs in space" ones. I've been playing this game for so long that I'm curious how exactly the final version of it was beaten into shape over time, and how different it feels to play over its development period.
You can be a central and large contributor on a project and know it. If you lord over other people with that knowledge and position and flex to stroke your ego, that's pride and arrogance.