1,188 karma · joined March 2, 2018
ALSO - something I think about a lot - if a all/most of the HumanLayer SaaS backend was open source, would that change your thinking?
concurrency abstractions keep changing (still transitioning / straddling sync+threads vs. asyncio) - this makes performance eng really hard
package management somehow less mature than JS - pip been around way longer than npm but JS got yarn/lockfiles before python got poetry
the types are fake (also true of typescript, I think this one is a wash)
the types are fake and newer. typing+pydantic is kinda bulky vs. TS having really strong native language support (even if only at compile time)
virtual environments!?! cmon how have we not solved this yet
wtf is a miniconda
VSCode has incredible TS support out of the box, python is via a community plugin, and not as many language server features
1. fire the async request, 2. store the current context window somewhere, 3. catch a webhook, 4. map it back to the original agent/context, 5. append the webhook response to the context window, 6. resume execution with the updated context window.
I have some ideas but I'll save that one for another post :) Thanks again for reading!
lemme know if/when you build the CLI agent that write my gherkin too
I'm thinking for a single particular application under test and a mostly-static group of SMEs who might be involved to respond/tune
The idea for incorporating feedback into the knowledge base is still coming together, in the prototype, LLM can classify the response as approval or not, and then if it's a rejection, llm will try to distill out facts/ideas from the response, e.g. "BigCorp and Acme.com are also using XYZ product" or "to learn more about pricing, you can book a meeting at LINK".
In the prototype, it then did a function call to add those as small chucks to the vector store, but you could also orchestrate that transparently if you didn't want to rely on the LLM reliably calling an `add_to_knowledge_base` function.
Longer term, I like the idea that I first heard of in BabyAGI, which is to store the messages leading up to an approval + the approval result in a vector DB, and use those historical approvals to derive up a confidence score for whether a particular action will be approved.
That stuff's more whiteboard stage than in code yet but I think it could be built.