226 karma · joined June 4, 2023
Where it gets interesting, is that we can save the execution plan that the big model comes up with and run with ONLY Moondream if the plan is specific enough. Then switch back out to the big model if some action path requires adjustment. This means we can run repeated tests much more efficiently and consistently.
I've always been skeptical of prompting techniques that ask the LLM to output a score or a confidence level numerically. The results from this experiment suggest that LLMs tend to understate their own confidence, and that "self-rated" scores prompted from LLMs may be generated more based on what the LLM thinks is a "safe" answer rather than an accurate representation.
The reason I'm curious about this area is because the startup I'm building does AI-powered E2E testing, and I'd like to more objectively figure out when a decision made by the agent is low-confidence so that it can be re-assessed.