Computer Graphics from Scratch
gabrielgambetta.com
gabrielgambetta.com
Long story short, today I’m incredibly grateful and excited to tell you that CGFS is coming out as a real book, with pages and all [2]. The folks at NSP graciously agreed to let me publish the updated contents, the product of almost two years of hard editing and proofreading work, for free on my website [3]. But if you’d like to preorder the printed or ebook version, you can use the coupon code MAKE3DMAGIC to get a 35% discount at https://nostarch.com/computer-graphics-scratch.
I’m still somewhat in disbelief that my work is getting published as a book, and this genuinely wouldn’t have happened without your support. So once again, THANK YOU :)
[0] https://news.ycombinator.com/item?id=19584921
[2] https://nostarch.com/computer-graphics-scratch
[3] http://gabrielgambetta.com/computer-graphics-from-scratch
And finding your Imdb profile was a nice addition on top ;)
All the best.
On modern machines, a pixel is just a set of three numbers indicating how much red, blue and green light should be shown at a particular point. Ex: {0.75 red, 0.0 blue, 0.5 green} for a kinda-dark, orange pixel. The GPU keeps a big 2D grid of these number-triples in memory and on a regular schedule sends out a copy over the DVI cable to your monitor. The monitor has a bit of memory to hold it's copy. And, it has hardware to scan over the grid of numbers to produce a sequence of voltage levels that are used to change the color of the points on the LCD.
There's a bit of math involved in how to do a good job representing colors with numbers and how to convert those numbers to voltages. But, at the most basic level, an image is just a big 2D grid of numbers. If you want to change the image, poke the grid. People want to change images a whole lot. So, we've developed pretty sophisticated hardware and software around poking 2D grid... But, that's a whole other topic.
The idea of a pixel is that it's the smallest area of the screen that can be independently controlled, thus the designation "pixel" for "picture element".
In most modern phone screens or monitors, each pixel is formed by a group of three smaller elements with fixed colors (usually red, green, and blue) but _variable brightness_. By controlling the brightness of these sub-elements, we can control what the overall color of the pixel appears to be - once you get more than an inch or two away from the screen, the light from the element group blends into what we see as a single color - so 100% green + 100% red + 0% blue looks like bright yellow. 50% each for red, green, and blue looks like a middling grey. You can usually see the structure of the pixel with a magnifying glass, though this is easier with an old TV or monitor than a modern phone.
These pixels are laid out on a regular rectangular grid, and your display controller will offer some way to set the color of each pixel and to then update them all at some (usually) regular interval, for example 60 times per second. In computers, it's common to keep a "frame buffer" around that stores separate values for the red, green, and blue components of every pixel on the screen. If these are 8-bit values, that implies that each R/G/B component can have 256 different levels of brightness, and in combination they allow each pixel to take on one of about 16,000,000 possible colors. So, a program can change pixel colors by writing different values into this buffer and waiting for the updated buffer to be processed by the display controller to change what the display is showing.
Of course, the electronics, physics, chemistry, timing, and logic of display generation have changed quite a bit since Pong. And not all displays even have pixels - vector displays used to be a thing, and are still used in some very specialized applications.
If you can zoom in enough to see subpixel elements, and pull up a color wheel, it's very intuitive to see that the screen is made of pixels, and the pixels are made of three elements.
https://en.wikipedia.org/wiki/RGB_color_model#/media/File:Ad...
The Pico-8 might be a approachable way to explore retrocomputing. It's an in-browser console modelled after GameBoy-era consoles.
https://news.ycombinator.com/item?id=18240375
I don't know about books, though. I remember reading the binary/computers chapter of How Things Work and being thoroughly confused at that age. There was some extended allegory about white mammoths and black mammoths.
Now I have both editions and have looked through them side-by-side. It's a bit unfortunate that some pages were dropped to make room for the expanded digital section in the newer edition, but well worth the trade-off for more/better explanations in the digital realm.
Just took a look over my bookshelf, but they haven't made it with me through moves.
I do still have Incredible Cross-Sections. That was another of my favorites to flip through. It looks like there are reprints and new versions. Would definitely recommend for kids.
https://www.amazon.com/Stephen-Biestys-Incredible-Cross-Sect...
Details like how and why pixels work are much easier to understand and retain if they're in some kind of context. Gamers learn about pixels to better understand the games or their gear. Devs learn about them in order to control them. Artists learn about them to better understand why computers mangle their art. So maybe choose an area the grandkids are engaged with and look for something there that also touches on display tech.
It's not like modern display panels and HDMI signals and digital protocols are simple, but you don't have to go into the details to understand how pixels and framebuffers work.
If you're trying to explain how pixels work, going back to 1980s era technology is a detour into arcane technical details that aren't really relevant any more. They're pretty cool to read about if you're into retro computer technology but they don't have much educational value or relevance to the modern day.
There's also a linear algebra appendix [0] that presents the operations, explains how to use them, and how they can be interpreted, without going in any theoretical depth about why these things are the way they are.
[0] https://gabrielgambetta.com/computer-graphics-from-scratch/A...
is that why you ended up working for improbable? ;)
You're not far off the mark, though! It was my other series of articles about client-side prediction [0] that brought me to Improbable's attention :)
[0] https://gabrielgambetta.com/client-server-game-architecture....
I'm sure you don't remember me, but we used to work together at Improbable – you sent me such a wonderful email when I left to go back to Sweden, thank you!
Anyway, just wanted to say congrats and hope you're doing great buddy. All the best! :o)
That's right, I was in Florida visiting family. Would love to reconnect, don't hesitate to ping me an email at marcus@madebystade.com if the feeling is mutual! :o)
What matters is I got the opportunity to reconnect with an old pal, I'll gladly pay the spam tax if it comes to that. Worst case scenario I guess I'm getting a new email address eventually. :o)
Good tip though, thanks!
My favorite chapters are those about drawing raster lines and triangles. One of my first programs was writing (in 386 16-bit asm) the routine to draw a triangle, and I love these algorithms dearly.
An important and beautiful thing that is missing in the book is the drawing of anti-aliased lines (easy) and triangles (not trivial!). I find that rendering carefully shaded smooth objects with a pixelized boundary loses a big part of the magic.
I explain a bit of this in the introduction, most of the time I present not the best/fastest algorithm, but the one that it's the simplest to understand. For example, I don't even mention Bresenham's algorithm; instead I present a super simple but inefficient one, which achieves two critical things: (1) you're drawing lines in no time, (2) it motivates the linear interpolation method that is then used for shading, z-buffering and texturing.
I think this is the final kick I needed, so thanks!
Things got very hard with ray tracing and later u,v texture mapping, etc. Making a note here so I can try your book out sometime!
Almost every modern source however uses engines or big frameworks and skips the parts I actually want to look into. Compiling a reading list or figuring things out myself of course would have been an option, but I guess others can relate to the problem of having too little time for hobby projects :)
Big bonus points also for trying to remain programming language agnostic! I immediately preordered your book and find it hard to express my excitement, can't wait.
It's a classic in the computer graphics field, like Knuth is for algorithms. I'd recommend it alongside the OP's new book.
EDIT: But for anyone who is reading this wanting to learn about practical computer graphics, CG:P&P is NOT the resource to use! Learning the low-level rasterization algorithms used in computer graphics is an important thing to learn at some point, just like learning assembly language provides valuable insights even if you never touch a line of assembler again. But if you actually want to write graphics code on modern hardware with GPUs, I'd highly recommend Real Time Rendering instead: https://www.amazon.com/Real-Time-Rendering-Fourth-Tomas-Aken...
What would you say are the prerequisites for this one?
Also from the description it sounds as if it also goes a bit more into 2D graphics too, is that correct?
You might think this has nothing to do with 3D graphics, but in fact there is a big overlap. The graphics driver + GPU takes a description of a 3D scene and projects it into 2D primitives whichs are rasterized onto a bitmap (the display) just as one might render a font glyph.
So IIRC the book begins with the basic theory and techniques of 2D graphics, and then shows how projective geometry can be used to render 3D scenes onto 2D displays, and progressively adds various corrections which we take for granted, such as perspective-correct interpolation of texture values, which these days is done automatically by the hardware.
I don't think there's any pre-requisites other than a basic understanding of beginner computer science, coding, and algorithms. If you can read Knuth, you can read CG:P&P.
Also, see my edit to the grandparent comment.
[0] https://www.amazon.com/Computer-Graphics-Scratch-Gabriel-Gam...
[1] https://www.barnesandnoble.com/w/computer-graphics-from-scra...
[2] https://bookshop.org/books/computer-graphics-from-scratch/97...
[3] https://www.penguinrandomhouse.com/books/645952/computer-gra...
This is one of the few tutorials that covers raster rendering. Lots of introductory material on ray-tracing, not a lot on raster graphics.
Where can students look for more stuff on rasterization after they finish this book?
Except, back then "from scratch" meant soldering together wires and transistors and such.
The big prize was displaying a vector Starship Enterprise on an oscilloscope.
Also 2019: https://news.ycombinator.com/item?id=19584921
and a bit here: https://news.ycombinator.com/item?id=19096322
There are several ways to approach this subject. This is a bottom-up approach - how do we get pixels on screen?
Other approaches start from "how do we represent a scene", in the sense of a scene graph or at least lists of vertices, triangle indices, and textures. That cuts the problem into "represent scene" and "render scene". That's a practical division, because that's the interface between what you put into a renderer and what the renderer does with it. Today, you're either using someone else's rendering system, or you're building the rendering system. Probably not both, except as an exercise.
This division is actually a bit above the OpenGL/Vulkan level. Vulkan is complicated because it's mostly about setting up the GPU to run a rendering pipeline, talk to displays, and other housekeeping. And memory management. If you have a library to manage that part, it's not so bad.
It's like old computer books, where you started out learning what the arithmetic/logic unit (the ALU) did, how it interacted with memory, what the instruction decoder did, and so on. This prepared you for assembly language programming. Few courses start out there today.
Incidentally, the caption for figure 12-1 is missing some symbols. It reads "Using instead of doesn’t produce the results we expect."
Maybe this is a really dumb question, but as someone who really knows nothing about computer graphics, how can you actually run these examples?
I realize that I could Google and research and find a list of "canvas drawing tools" in a variety of languages that I could then evaluate (despite, again, knowing nothing about graphics), but it'd be SUPER-SUPER-awesome to have a quick note "there are lots of ways to go about making these things happen in the real world, but here's one I recommend if you have no other priors." :-)
Really cool book idea, keen to have a go at it!
Python has PyGame, or you could just plot to PNGs.
Java awt has a canvas.
C/C++ has SDL/SFML or also plotting to PNGs (or if you want to go wild, embedded GPUs are an option).
You could also do all of this in BASIC if you write the canvas commands. All of this is doable on a Commodore in software albeit slowly.
This might be the easiest way - make a trivial HTML page with a <canvas> tag and start setting pixel colors in it :)
I'd also highly recommend Ray Tracing in One Weekend[1], which starts you out with literally just a c++ program to dump bytes into a file which can be opened with most standard image/document viewers. Then you don't have to think about or set up any shaders before you can start to experiment with the rendering algorithms.
[1] https://raytracing.github.io/books/RayTracingInOneWeekend.ht...
Edit: formatting.
I've been teaching myself 3D/CG programming in my spare time using free web resources, I've been trying to learn it "the hard way" by doing everything relatively from scratch, instead of using a high-level engine I'm using low-level libraries writing my own shaders and implementing my own scene graph etc. This book really seems right up my alley and looks like it'll teach me a lot.
For anyone else similarly interested in learning to do CG from scratch (particularly on the web), some other great resources I've found are: https://thebookofshaders.com/ https://webglfundamentals.org/
I've took the liberty of rewriting the 'Perspective projection' demo (1) in more up-to-date Javascript, using ES6 classes and all the other niceties that modern browsers have. I also used the regular canvas line drawing methods to focus on just the vertex-to-line conversion, and updated the code so that it will work on any resolution. https://codepen.io/hay/pen/gOLazpm
1: https://gabrielgambetta.com/computer-graphics-from-scratch/d...
As someone who works with CAD/CNC machines on a daily basis, this is torment. XY is the floor and XZ, YZ are "walls". Same would go for a 3d printer. They're not that arbitrary.
I think the choice of 3D up axis is related to the expected look direction. 3D printer control software evolved from CNC control software, which is looked at from top down, so up is on the build plate plane. Video game commonly looked at from the side, so up is up. Up is then often associated with Y.
And then there the argument about right hand vs left hand coord system...
That was the opposite of fun when I was noodling with writing an OpenGL-based rendering engine for a "fly a spaceship around a 3D maze" game, decade and a half ago.
As to why I prefer right-handed? The lin-alg textbook I had at uni uses a right-handed coordinate system for what it calls "the perceivable room" (yes, it is "the book with the cow on", for those in the know).
I don't know how this came to be. You make a very good point that XY as the floor makes a lot of sense, especially if you're coming from the context of "I'm drawing a floorplan on a horizontal table" - the floorplan matches the orientation of the actual floor so XY is the most natural choice.
I suppose for a similar reason CG adopted the "XY is a wall" convention, since you start with XY on the screen (a vertical plane in front of you) and only later add depth?
I've been reading the book and enjoying it so far. Thank you.
The numbers on touch-tone phones (as opposed to rotary-dial kids) and the numbers on computer keyboard number pads are reversed:
1 2 3
4 5 6
7 8 9
vs 7 8 9
4 5 6
1 2 3
Why? No one knows...[1] https://gafferongames.com/
[2] https://gafferongames.com/post/reliability_ordering_and_cong...
[3] https://gafferongames.com/post/reliable_ordered_messages/
[0] https://gabrielgambetta.com/client-server-game-architecture....
I see you reference Gaffer on Games and the Valve Latency Compensation article, both are great for understanding game networking and helped me greatly in shipped titles especially prior to Unity/Unreal using stuff like enet [2]/RakNet [3]/custom. Lots of networking libs inherited the best from those like reliable UDP, channels and dealing with NAT. Both are excellent libraries to help understand the full picture and get started. Lots of game engines networking libs are based on those.
[1] https://developer.valvesoftware.com/wiki/Latency_Compensatin...
[3] http://www.jenkinssoftware.com/ + https://github.com/facebookarchive/RakNet
I got a good chuckle out of your technical reviewer: "Alejandro currently works in the GPU Software group at a leading consumer electronics com- pany based in Cupertino, California." I wonder where he works? ;)
Having a single course written by a single person that presents both major rendering techniques is amazing.
[1]: https://web.archive.org/web/20120212184251/http://devmaster....
Also the original 1988 paper by Pineda is a surprisingly easy read, and it's only 4 pages![1]
[0]: https://fgiesen.wordpress.com/2011/07/06/a-trip-through-the-...
[1]: https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.15...
And of course not to forget all the other approaches to modeling geometry like implicit surfaces, meta-balls, signed distance fields, constructive solid geometry, point clouds, splats, voxels, etc ...
Constructive solid geometry I do mention as an extension to the raytracer [0], but it doesn't have a dedicated chapter. Perhaps for the 2nd Edition? :P
[0] https://gabrielgambetta.com/computer-graphics-from-scratch/0...
I did some research a year ago. Found a few alternatives, yet none of them was good enough. Here's the issues I remember.
1. It's complicated. Cubic segments can self-intersect or contain singularities.
2. Stroked curves are used a lot. To build stroke edges from the curve, need to offset the curve by half width of the stroke. When you offset a polyline you gonna get another polyline. Yet the offset of Bezier splines is not generally representable as another Bezier spline.
3. In some use cases, hardware-implemented MSAA is the best way of AA. Polygon pipeline gives you that for free. Producing SV_Coverage in pixel shaders to achieve same effect for mid.points of triangle is hard and inefficient.
4. Games have been pushing GPUs for high polygon count for couple decades now, GPUs became really good at that. Also GPUs have early Z rejection that drops pixels or even larger blocks before pixel shader stage. You can't have that if you sending large primitives and computing curves in pixel shaders. Polygon pipeline is not necessarily slower.
It stops being trivial if you want these circles to look nice and render fast.
Take a look: https://github.com/Const-me/Vrmac#vector-graphics-engine
https://gabrielgambetta.com/computer-graphics-from-scratch/d...
if you have not checked it out yet, try it. it is quite enjoyable.
Vulkan and DX12 however are the work of the devil.
I initially tried using D3D/Win32 APIs to perform the high-level drawing operations for me, but found that for the scope of functionality that I required, these interfaces were far too heavy-handed. These also have lots of platform-specific requirements and mountains of arcane frustrations to fight off.
I didn't need to interface with some complex geometry or shader pipeline in hardware. Raytracing is hilariously out of scope. I really just needed a dead simple way to draw basic primitives to a 2d array of RGB bytes and then get that to the display device as quickly as possible. What I ended up with is something that isnt capable of very much, but can run on any platform and without a dedicated GPU. I also feel like this was a much better learning experience than if I had slammed my head into the D3D/opengl/et. al. wall.