I've been using DuckDB similarly and loving the simplicity. What sort of latency are you seeing with your setup? I'm hitting 300ms to query against 10M vectors and could see that becoming a bottleneck if going into the hundreds of millions.
[1] https://blog.pgvecto.rs/vectorchord-store-400k-vectors-for-1...
Depending on how the wrapper is implemented you can get different numbers. With the raw libraries, on larger machines, I'd expect 200'000 requests per second for 1 Billion entries.