Nvidia released Nemotron 3 nano recently and I think it fits your requirements for an OSS model:
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B...
It's extremely fast on good hardware, quite smart, and can support up to 1m context with reasonable accuracy