Do you use a tool for this? Is there some sort of tool which collects evals from live inferences (especially those which fail)
This is a use of Rerun that I haven't seen before!
This is pretty fascinating!!!
Typically people use Rerun to visualize robotics data - if I'm following along correctly... what's fascinating here is that Adam for his master's thesis is using Rerun to visualize Agent (like ... software / LLM Agent) state.
Interesting use of Rerun!
darin@mcptesting.com
(gist: evals as a service)