They’ve never said that they hand–wrote any assembly, though I suppose it’s not impossible. Mostly they were simply diligent and focused on never regressing performance, and on continuously improving it. They even changed game mechanics over the early access period to improve performance.
For example, originally a conveyor belt applied a speed to any object in the same tile. Each frame the position of the object would be updated based on the speed, just the same as any other moving object. You could place down up to three objects side–by–side on the belt, and they would all travel together. But eventually they realized that this is a waste of time, because inserters only ever place items in specific positions on the belt, and the items are all the same size. This means that they could instead just keep a bitmask of which spots are filled, and a queue of items that will pop out of the edge of the tile. Even that is wasteful though, because you might have 1000 tiles of belts all in a row, so you can group them all into a single queue. Now you only have to put new items into the beginning of a single queue and take them out of the end, instead of doing that for 1000 queues. On the other hand, this means that you can no longer put down items at arbitrary positions on a belt; they will now snap into one of the 12 slots in one of the two lanes.
A similar thing happened to inserters. It used to be that they had a certain amount of rotational speed, and would calculate a new position and rotation for each of their three two joints every frame. Eventually they realized that 99% of the time they would be in the same few positions over and over again; it is only when they are chasing an object on a belt that is too fast that they need to calculate a position. The rest of the time they just loop through a fixed animation. This gave Wube the opportunity to tweak the animation cycles so that they are exactly the same length in all orientations of the inserter, which is definitely a good thing.
And of course they consistently paid close attention to data locality, so that every time they have to work on a collection of objects they maximize the amount of useful data that fits into the CPU cache. I haven’t seen the code, of course, but I assume that this means that they are using an ECS–style system where you have a dense array of just the position and speed of each biter, or just the animation frame for each inserter. The code can then loop over those dense arrays knowing that it doesn’t have to test the type of each item in the array (few or no branches needed), and that every byte loaded into cache will be used during the loop (maximizing use of the memory bandwidth).