What I Learned from Watching Notch Code
gun.io
gun.io
For the record, this is completely contrary to his testing method for Minecraft. Every time there is a new version of Minecraft, there are more bugs than features. This is even true when a "bugfix" release comes out. More bugs are introduced in a "bugfix" release than are fixed. Seriously, check it out: http://www.minecraftwiki.net/wiki/Version_history
I say this as the owner of a library to provide bot API to Minecraft. https://github.com/superjoe30/mineflayer
All I'm saying is... "thoroughness" and "testing" are not two words I expected to be in the same sentence as Notch's methodology.
Also, while that is the most recent glaring bug, in the past there have been multiple occasions where a release broke trees so that the leaves didn't decay after the trunk was removed. Punching down a tree is something you do in the first 30 seconds of a normal minecraft game.
But that’s not really the point. There is no way to test Minecraft thoroughly if you are pretty much on your own. It has too many features. Minecraft is too complex for this method to work.
Not that it matters, really.
Woah, man. Woah.
I think you're missing the primary benefit here, which is that he avoids regressions. By the time you are adding feature 256, you have 255 features which are going to possibly be affected by the code changes you introduce. This is a huge part of software engineering, and one that has gotten lots of attention over the years.
The process you describe for your development sounds extremely tedious, but probably won't break down until month 2 of development on a team of one. Once you reach a level of complexity beyond this, that's where automated testing proves it's salt.
The style of testing you're describing is commonly referred to as integration testing or acceptance testing, because it is designed to test the full stack in harmony. There are a number of great frameworks out there to help you do this. Cucumber is the one that's gotten the most love in the circles I am familiar with. You can write your steps in python or javascript, so don't worry that it's written in Ruby.
The typical thing to do once you've started doing automated testing is to actually write your tests first, watch them fail, then write the code to make the test pass. This forces you to ensure you have good test coverage (every feature is tested) and has been shown to result in better designed systems.
You have tons of reading to do if you want to learn more about this, but hit me up if you want a basic rundown.
And it tests the tests too!
To expand on testing a bit: at one end, there's unit testing where one tests maybe a single object's behavior in the game. I think for game dev, testing for absolute values are not as important as deterministic ones. We would rather want to know if a player has jumped, died or shot something.
Eventually, a functional testing is needed when objects interact with one another. It's like a system of systems. That one's more complicated to design. It might go as far as designing a test that runs through a complete level within miliseconds. Anyone have ideas/resources on this?
For the server and client both, I really like Jasmine as a testing framework. Of course, there's always cucumber :)
* possibly an exaggeration, but it's definitely very very bad.
And since there was only 1 working communication port, how did I debug? How did I even run a test?
The board I was developing on was amazing. It had 1 (one) LED.
Blink.
Blink.
Blink (fast).
As a bonus, it used a dodgy proprietary compiler with a whole sack of undocumented "technical limitations."
The good part is I now have a really good understanding of pointer manipulation. The bad part is using it sometimes makes me twitch.
Edit: Here's someone who's got us both beat: http://web.archive.org/web/20070613032334/http://ipodlinux.o...
But I guess then you'd have to debug that code and you'd be back where you started.
At the company I work for we rely mostly on manual testing for testing the games and applications we make. We do some automated testing in the form of Tinytask tasks that are left on overnight to hammer an application. In terms of release testing, especially with a custom game engine, there's no easy way to codify a test which actually plays through a game where there is some random mechanic, or checks for things like the visual accuracy, windows flashing up on the screen when they're not supposed to, etc.
Web development frameworks like Selenium are great for UI testing, but they require identifying interface elements by their IDs, or performing a 100% image recognition match for targeting. If anyone knows of a good UI testing framework that would work for games/DirectX apps, I'd definitely love to hear about it.
I'm not bagging automated testing, I think it's a great way to save time and avoid the boring bits. I just don't think it's very relevant to game design, but for the OP's enterprise app job it would definitely be suitable.
Of course, if the game does need heavy-duty serialization of everything, all bugs are potentially deadly. And this is largely the case for the line-of-business, social, and productivity apps, because that data is considered mission-critical. Corruption is not OK, reverting to backups is to be avoided. When a game has that kind of requirement, things get a lot tougher - and so games have naturally evolved to favor minimizing the save data to nothing or a few stats.
And this is borne out by looking for contrapositives. There are two genres that have a history of tending to be buggy because of some form of long-term data corruption: Large-scope RPGs and turn-based strategy games. It's hard to reach the bugs in those games, so it's also hard to fix them.
Because all of these apps are based on a game engine, we get the complex and hard to test bugs that games get (texture cache corruption, null pointers in the scene tree, etc.). There's much more serialization and pressure than with pure games (but still not as much as with enterprise applications) because we're dealing with transactions with 3rd party services, or because a screw up could corrupt a user's whole music collection, or because users don't expect their music player to crash every 4 hours.
The game engine adds a huge amount of extra variability to our applications; we have to not only watch out for obscure bugs in our custom script code (which the engine silently tries/fails to execute anyway), but also in the (C/C++ based) game engine it is driving. The upside to using an engine with a simple scripting language is the shorter development times to get things off the ground and the high-performance/shiny visuals, but I feel it costs us much more down the line in terms of stability and extensibility.
Edit: It's both scary and awesome that you rigged a game engine to manage a music library a-la iTunes. Are you profitable?
http://www.unlimitedrealities.com/blog/video-the-dell-stage-...
Those issues I mentioned are examples of things that we have actually run into.
edit: Our specialization in touchscreen development is really what drives the company, but hardware accelerated graphics are also a huge draw.
Yea I think lag in touchscreen interaction is a huge problem. I think this is largely a result of hardware. We've had to ship on atom-based hardware with 2003-era DX9 graphics. Even the latest Intel chips with DX10 are horridly slow when it comes to graphics.
The other side is actual touchscreen hardware. The machine used in the video is an HP Touchsmart, which uses an optical touchscreen panel. These panels have a latency of around 100-200ms, and that's before we even start processing the touch event information. Capacitive touch sensors are much better, but the're expensive to manufacture above about 10 inches.
In the end lag/accuracy is a reality of the low-cost hardware OEMs use, and there will always be some trade-off between hardware cost and performance.
If you're feeding input into a non-deterministic game the character won't be exactly where they were last time, etc. You could keep track of the movement offset and see if it's within an acceptable range, or you could check the character's speed at two times and make sure they're accelerating properly, etc. Design a test level to highlight the potential problems (weird ground polys, whatever) and write an in-game script to test the behavior.
I think I could even script visual checks - mostly. You'd make, for example, a level prone to Z-buffering errors, where the colors were chosen to make it obvious - like a red wall showing through a white one. Save a stream of screenshots and compare them. Trivially, check for red. More complex, and perhaps better, check for high-contrast areas. It wouldn't tell you if the picture looked good overall but it'd be fairly good at finding instances of that one bug.
And for unit tests, don't be so quick to write such a good test that throwing it away is painful when you want to refactor the code.
tell application "BBEdit"
save front document
end tell
tell application "Terminal"
if (count of windows) is 0 then
do script "/path.to.my/make.command"
else
do script "/path.to.my/make.command" in window 1
end if
end tell
tell application "Google Chrome"
activate
end tell
My make.command (simplified): #!/bin/sh
coffee -c mycode.coffee
lessc -x mycode.less > mycode.css
rsync -avz --delete --force --exclude ".DS*"
-e "ssh my.ppk" /localpath/www user@me.com:/remotepath/www
On my server, I run a changed.php file: function fileHash($fn) {
echo sprintf("%u.", crc32(file_get_contents(dirname(__FILE__) . $fn)));
}
header('Access-Control-Allow-Origin: *');
fileHash('mycode.js');
fileHash('mycode.css');
On all my HTML pages, I run this piece of CoffeeScript (compiled above to JS): # defined elsewhhere:
# debugmode (bool)
# reloadIfChangedLast (empty string)
# getCurrentView() returns #pageHashURL
reloadIfChanged = ->
if not debugmode
return false
$.ajax
type: 'GET'
url: 'me.com/changed.php?rnd=' + Math.random()
error: ->
setTimeout reloadIfChanged, 500
return
success: (data) ->
if reloadIfChangedLast and data and
reloadIfChangedLast isnt data
reloadIfChangedLast = data
document.location.href = 'index.html?rnd=' +
Math.random() +
'#' + getCurrentView()
else
if data
reloadIfChangedLast = data
setTimeout reloadIfChanged, 500
return
return
It may seem like a lot of hassle but for me it's just copy-paste and edit a few things per project. And the benefits are tremendous.All of Apple's official applications support scripting, so do most good third party apps. You don't need extra JS on your page to autorefresh Safari. This little snipped reloads the front-most Safari tab:
tell application "Safari"
set sameURL to URL of tab 1 of front window
set URL of tab 1 of front window to sameURL
end tell
You can even send Safari a snippet of JS to run. It's possible to automate any UI interaction by sending mouse clicks or keyboard events to tell application "System Events".You can look at any app's AppleScript dictionary using the AppleScript Editor. Go File -> Open AppleScript Dictionary ... and then pick your app.
He began by building the engine, and to do this he used
the ‘HotSwap’ functionality of the Java JVM 1.4.2, which
continuously updates the running code when it detects
that a class has changed.
If anybody at Google is reading this, if you add this
feature to the Android emulator and I will literally
drive to your house and kiss you on the mouth.
Indeed, the fact that Android doesn't already have this makes it hard for me to take it seriously as a platform. Immediate feedback is crucial to any development environment.Edit: I know 30s is still nowhere near ideal. Hell, on Linux Chromium (a massive project) we have faster iterations than that. But it's good compared to the Android emulator.
I think you should be able to go even faster, but I need to dig into how Ant actually works. It seems to be redoing more work than necessary for each build.
Notch was able to take advantage of immediate feedback for the most part of his coding. The first hours spent on the rendering was tested by watching the world being rendered live. I've been doing the same developing a game in Clojure. When he got as far as the gameplay, he slowed down considerably, since he had to wander around in the game world to test each new feature. He for example made temporary shortcut passages, so he was able to test the boss monsters, without actually having to pass through the levels.
He is a person who can concentrate on delivering results and hack away at code for long stretches. This is the kind of code he presumably has done again and again for years. Who professional web developer can't hack together a small site as quickly? Where most people fail is attention span and drive! I know my unfinished projects speak of that :)
Most interesting part to me, of watching Notch code, was to see how he used his tools. And it was inspiring and motivating to see the progress. The actual code was very hackish but quality code is not important in a throw-away project anyway :)
For those who didn't RTFA, though because this isn't Slashdot there shouldn't be any, Notch is reputed to have given some segment of his game a complete replay every time he made a change though it isn't mentioned exactly how often this is or how big the change is. A note is made that his build scripts make this almost instantaneous.
This is a problem because manual testing takes forever and causes tester fatigue. You stop doing a good job.
It can be harder to test behavior in a 3D game than a text filtering app but imho a good design is a testable one. (This does not necessarily mean I'm that good of a designer yet...)
Automate, automate, automate.