It uses the rasterization HW for large triangles and uses a software rasterizer in compute for small triangles.
The HW rasterizer is not that efficient if the triangles are tiny, which is the case for nanite.
But that is one generation of GPUs away when they will make some computational shortcut for small triangles, which to me seems to be rather trivial to implement.
Not really trivial if you have to support any input set of triangles and don't know much about them. Software rasterizion exploits things like localized chunks of triangles, but the hardware rasterizer does not know about that in advance. Also, these kinds of software rasterization algorithms add limitations, which aren't much of an issue with your own rendering pipeline that specifically knows how to deal with these limitations and works with them.
For small triangles, they use a software rasterizer into the V-buffer. Obviously for 1px triangles since they don't want to waste quad overdraw, but I think they found it's faster up to 12px/tri or so on AMD.
Part of Nanite is software rasterization by rendering triangles with 64 bit atomics. You can simply draw the closest fragment of a triangle to screen via atomicMin(framebuffer[pixelID], (depth << 32) | triangleData).