Six years in software engineering at early-stage startups, mainly on production TypeScript (React, Next.js, Node), with two businesses of my own along the way. I now work at the intersection of agentic coding and evaluation. I built an agentic coding tool by hand to understand agents below the frameworks, then built a lab to evaluate them, and published both the study and the reasoning that made me stop. Currently working on a floating persistent terminal and running an independent audit of the session_success metric in the SWE-chat dataset. Open to roles building or evaluating agent systems.