It's a bit jarring that, when you run the command twice in the same directory, you get a different root hash. IIUC, this nondeterminism comes from two sources: 1) each file is hashed with a random nonce, so as to distinguish files with identical contents; and 2) the work is distributed among goroutines on a first-come-first-serve basis, so the position of each file in the tree varies from run to run.
The latter shouldn't be too difficult to solve; just match each file's hash to its index from when the directory was walked. You can then use that index as a nonce, resolving the other source of nondeterminism.
I understand that you can still build and validate proofs despite the nondeterminism, but there's a good reason to eliminate it: doing so also eliminates the need to store the tree nodes! That is, storing the nodes becomes a performance optimization, rather than a necessity. (Oh yeah, another bit of UX feedback: I made the mistake of outputting the tree file to the same directory I was hashing, inadvertently introducing another source of nondeterminism! Maybe print a warning or something if the user does that.)
Lastly, I suggest representing your Merkle tree as a BNT stack (https://eprint.iacr.org/2021/038), as this will allow you to compute Merkle roots and build/verify proofs in streaming fashion, rather than allocating a big slice to hold all of the leaf hashes. The blake3 package you import uses this approach ;)