Our gating model(the heart) in Humiris' system functions as an intelligent gating mechanism that dynamically selects and orchestrates multiple large language models (LLMs) based on predefined parameters such as cost, performance, privacy, and speed. While the exact implementation specifics can vary, here’s how it might work under the hood:
The gating model evaluates each query and the available models' characteristics to determine the optimal routing. This involves dynamically assigning weights to multiple parameters and scoring models based on the query's requirements. The architecture combines elements of both machine learning and decision-making algorithms rather than relying solely on traditional decision trees or reinforcement learning.
Neural Networks with Softmax Activation:
A neural network trained to route queries based on encoded query features and user priorities. The softmax function outputs probabilities for each model, and the model with the highest probability is selected (or multiple models in collaborative tasks).
Reinforcement Learning (RL):
In advanced systems, reinforcement learning may be employed. The routing model learns optimal routing strategies by maximizing rewards (e.g., high response quality, low latency, reduced cost). RL can also adapt to new models or parameters over time, improving efficiency through trial and feedback loops.
The routing model in Humiris likely uses a hybrid approach, combining machine learning (neural networks) for dynamic decision-making with principles of multi-criteria optimization. The advanced mode incorporate reinforcement learning for adaptability in complex and evolving environments