HNHacker News
TopNewBestAskShowJobs

anerli

226 karma · joined June 4, 2023

submissionscomments
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Oh this is interesting. In our case we are being very specific about which types of prompts go where, so the planner essentially creates prompts that will be executed by Moondream, instead of trying to route prompts generally to the appropriate model. The types of requests that our planner agent vs Moondream can handle are fundamentally different for our use case.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
So it's key to still have a big model that is devising the overall strategy for executing the test case. Moondream on its own is pretty limited and can't handle complex queries. The planner gives very specific instructions to Moondream, which is just responsible for locating different targets on the screen. It's basically just the layer between the big LLM doing the actual "thinking" and grounding that to specific UI interactions.

Where it gets interesting, is that we can save the execution plan that the big model comes up with and run with ONLY Moondream if the plan is specific enough. Then switch back out to the big model if some action path requires adjustment. This means we can run repeated tests much more efficiently and consistently.

anerli··on Can LLMs accurately evaluate their own confidence?
I ran a simple experiment to try and understand whether self-rated answer confidence reflects the actual probability of the LLM generating that answer.

I've always been skeptical of prompting techniques that ask the LLM to output a score or a confidence level numerically. The results from this experiment suggest that LLMs tend to understate their own confidence, and that "self-rated" scores prompted from LLMs may be generated more based on what the LLM thinks is a "safe" answer rather than an accurate representation.

The reason I'm curious about this area is because the startup I'm building does AI-powered E2E testing, and I'd like to more objectively figure out when a decision made by the agent is low-confidence so that it can be re-assessed.

← PreviousPage 3 of 3