1,085 karma · joined September 2, 2019
The harness was extremely simple: A handful of MCPs + Skill.MDs and OpenCode with a stayalive daemon inserting "continue" every time it went idle
https://hyperengineering.bottlenecklabs.com/p/the-infinite-m...
To me, Python's best feature is the ability to quickly experiment without a second thought. Conda is nice since it keeps everything installed globally so I can just run `python` or iPython/Jupyter anywhere and know I won't have to reinstall everything every single time.
Surely we could do better? Testimony doesn't assuage my concerns that the process may not be tamper proof.
Every time I try to get to the bottom of this, it always boils down to "trust the system" which makes me uneasy.
For the experienced lot of us, I've heard many call it "hyper engineering"
I also fire off tons of parallel agents, and review is hands down the biggest bottleneck.
I built an OSS code review tool designed for reviewing parallel PRs, and way faster than looking at PRs on Github: https://github.com/areibman/bottleneck
However, fans of the game are calling it "Demastered".
Shockingly, the best code review tool I've ever used was Azure DevOps.
https://hegel-ai.com https://www.vellum.ai/ https://www.parea.ai http://baserun.ai https://www.traceloop.com https://www.trychatter.ai https://talc.ai https://langfuse.com https://humanloop.com https://uptrain.ai https://athina.ai https://relari.ai https://phospho.ai https://github.com/BerriAI/bettertest https://www.getzep.com https://hamming.ai https://github.com/DAGWorks-Inc/burr https://www.lmnr.ai https://keywordsai.co https://www.thefoundryai.com https://www.usesynth.ai https://www.vocera.ai https://coval.ai https://andonlabs.com https://lucidic.ai https://roark.ai https://dawn.so/ https://www.atla-ai.com https://www.hud.so https://www.thellmdatacompany.com/ https://casco.com https://www.confident-ai.com
I did a solo hackathon project [0] using poke-env[1] and found it to be pretty easy to get started with. Hoping for an easy-to-use API with this as well.
[0] https://x.com/ditzikow/status/1922004651790000521 [1] https://github.com/hsahovic/poke-env
Also very nice of them to include extensible tracing. The AgentOps integration is a nice touch to getting behind the scenes to understand how handoffs and tool calls are triggered
Given the size of the niche (developers building voice agents), do you find there's a lot of demand for testing and observability? From my anecdata, many of the voice AI agent builders are using SDKs and builder tools (Voiceflow, Vapi, Bland, Vocode, etc). Observability is usually already baked-in pattern with these SDKs (testing I'm not so sure of).
One conversation I had with a voice agent builder: "Our product is complex enough where external testing tools don't make sense. And we know when things are not working because we have close relationships with power users and companies." Whose problem are you solving?
Your tool looks very powerful, but might the broader opportunity be just to use your evals to roll out the best voice agents yourself?
Besides, there's a warning message for when you specify a model without a known tokenizer.
If you're upset with the implementation, you can always raise an issue or fix it yourself
At this moment, Tokencost uses the OpenAI tokenizer as a default tokenizer, but this would be a welcome PR!