Smaller models will likely not have 32k context windows.
And I think the optimized backends should implement that sliding 16k context soon...
Anyway, point is a huge context really helps certain types of queries, and VRAM usage is reasonable with a 7B model.