If you don't have something running somewhere, you don't have an agent, you don't have a harness. You've got a token generator, an LLM from the 2024 era.
You could build something with storage and composition out of MCP functions, but come on, have you seen how LLMs - particularly budget LLMs - try and invoke functions reliably? The amount of retries you have to hide, feedback you need to send back to the LLM about what it did wrong. Parameters get replaced with synonyms, arrays are passed for singular arguments and vice versa, structured inputs are flattened, etc.
So maybe you fine tune on interactions with your subset of MCPs, to improve reliability. But all you end up doing is reinventing a Unix-like command line, poorly.
Firecracker micro-VMs, gVisor, wasm sandboxes. There are ways to make this work that aren't heavyweight. Giving LLMs tools that they've seen how to use millions of times in training corpora just works better.