Now, you are right that this might not matter much if you only consider coins in particular. But everything in a level wants to have additional surface details. The ability to have normal maps everywhere was a huge part of the jump in visual quality a couple of console generations ago.
First, modern GPUs are capable of handling millions of triangles with ease. Games are almost never vertex bound. They are usually pixel bound or bound by locking. In other words, too many or too costly of pixel operations are in use or there is some poor coordination between the CPU and GPU where the CPU gets stuck waiting for the GPU or the GPU is underutilized at certain times and then over burdened other times.
Adding two maps adds two texture fetches and possibly some blending calculations. It's very hard to weigh the comparable impact because different situations have a huge impact on which code gets executed.
For a model with many triangles, the word case scenario is seeing the side of the model with the most faces. This will cause the most triangles to pass the winding test, occlusion tests, frustum tests, etc. This will be relatively the same regardless of how close the coin is to the camera.
For the normal mapped test, the worst case scenario is when the coin fills the view regardless of its orientation. This is because the triangles will then be rasterizied across the entire screen resulting in a coin material calculation and therefore the additional fetches for every pixel of the screen.
Also, when it comes to quality, normal maps have a tendency to break down at views of high angular incidence. This is because though the lighting behaves as expected, it becomes clear that parts of the model aren't sticking up and occluding the coin. This means such maps are a poor solution for things like walls on long hallways where the view is expected to be almost perpendicular to the wall surfaces most of the time.
There is a solution called Parallax Occlusion Mapping. Though it is expensive and I don't see it in a lot of games.
https://en.wikipedia.org/wiki/Parallax_occlusion_mapping
https://www.youtube.com/watch?v=0k5IatsFXHk
You'd have a bad time including all of that in the actual model.
As always, what is the right approach depends. If you want to run the game in a very memory constrained environment (early game consoles), the approach of using polygons might make more sense. But nowadays we have plenty of vram and if the entity is used a lot, 1 texture for 100s of entities makes it a worthwhile tradeoff.
Edit: thank you for the corrections, I'm not heavily into game programming/rendering but am familiar with 3D/animation software, but the internals of rendering is really not my specialty so I appreciate being corrected. Cunningham's Law to the rescue yet again :)
That hasn't been true for many years.
> as you'd use more vram when using textures
For the equivalent amount of detail, not really. Vertices are expensive, normal maps are encoded very compactly.
This is a game from 2007, 16 years ago (and its development would have started in 2005), so that it has not been true for many years is not a useful assertion even if true.
Or worse, it confirms that it used to be an issue, and thus may well have affected a game created “many years ago”.
And modern GPUs are plenty fast enough to do quite a bit of work per pixel, at least until the pixel count gets silly (resolutions of 4K and beyond), hence all the work going into AI upscaling systems.
(But the first coin image is from Super Mario Galaxy, a Wii game, and the Wii had a rather more limited GPU with a fixed-function pipeline rather than fully-programmable shaders)
If you were coming to it from the PS2 (like I was), it was super-freeing and amazing and so much more flexible and powerful than anything you were used to. But if you were coming to it from the Xbox (like many of my co-workers were) it was more like suddenly being handcuffed.
True for the CPU, but quite unfair to the Wii's ArtX GPU, which was substantially faster and more featureful than the ATI r128 in the PB G3.