- If you have a fixed time budget and increase the GPU memory+compute available, you can directly query a bigger model. Raw models are basically giant lookup functions, and without the extra memory+compute, they'll spill to slower layers of your memory hierarchy, e.g., GPU RAM -> CPU RAM -> disk. Likewise, with MoE models, there are multiple concurrent models being queried.
- Most 'good' LLM systems are not just direct model calls, but code-based agent frameworks on top that call code tools, analyze the results, and decide to edit+retry things. For example, if doing code generation, they may decide to run lint analysis & type checking on a generated output, and if issues, ask the LLM to try again. In Louie.AI, we will even generate database queries and run GPU analytics & visualizations in on-the-fly Python sandboxes. These systems will do backtracking etc retries, and > 50% of the quality can easily come from these layers: LLM leaderboards like HumanEval increasingly report both the raw model + what agent framework on top. All this adds up and can quickly become more expensive than the LLM. So better systems can enable more here too.