it's kind of trivial today.
the first layer could be a mix of nlp and zero-shot classification to clarify the nature of the request. Then using LLM deconstruct the request into several specific parts that would be sent to specialized LLMs. Then stitch it back together at the end again with LLM as the summarization machine.
Problem is running so many LLMs in parallel means you need quite a bunch of resources.