For occlusion culling it's a bit more tricky because you can either do it on the CPU in broadly two ways, either do low-res raycasting / software rendering like in the article on the CPU and cull based on that. This is an adaptive workload, you can give it more threads or CPU power and it scales for better culling which results in less pixels rendered on the GPU.
You can also use GPU culling but that's more complicated to do and that uses the GPU which creates a catch-22 - you want to use GPU culling to reduce GPU load but integrated GPUs don't cope well with compute shaders and memory bandwidth in general, so doing a culling pass might wipe out any culling gains you might have.
And dedicated GPUs have the raw power and memory bandwidth to just submit everything in your frustum and get most of it depth-rejected.
I wonder if it would be even faster to create a connectivity graph on the CPU, like each chunk knows whether a neighbour is visible and vice versa. On rendering the chunk graph is walked and the visible chunks are submitted, kind of like a primitive garbage collector to determine liveness. The culling would be worse but I presume traversing a fairly small (few thousand elements) list is quite a bit cheaper than rendering the "mipped" occlusion boxes, but do let me know if this is wrong.