I'm not sure how enforceable that is tbh.
18 karma · joined March 14, 2024
kid 01205799dbea47c29f8cdda92f7c19c46f64d98b4de612bf063de63fd43caca59eaf0a
I'm not sure how enforceable that is tbh.
Yes. Which is why it's in my app code, committed to source code, and with a guarantee of determinism. If you don't consult the runtime information (i.e. data), there's no reason why it has to be generated at runtime.
> We are working on making the execution more dynamic, possibly allowing code to be generated and executed progressively when data arrives and new prompts are given.
I think you just described how a react agent works.
> you cannot pause it to wait for human interaction for a day, then resume it
Sure you can! Have you seen langgraph, trigger, temporal?
Also, how do you pause based on an AST that might get invalidated if you recompute it on new contextual information?
Does the human keep approving until your AST works?
Sorry to be terse, but surely I'm missing something here.
Since you're computing the AST upfront, you'd get it wrong more times than a react agent that does it as and when needed. So, do I pay OpenAI everytime you make a mistake on the AST?
Or is your claim that your AST is infallible?
> do not even need to know all the data to perform operations on it or make decisions
If I know code-generation is going to be possible without any contextual information, I might as well generate the code using copilot or Curor and commit my code. Why do I need a runtime agent to do it?
What if the control flow has to change based on a result it receives?
What if the plan up front is wrong and needs to change halfway? Do I run the entire thing again with a new plan? What if my tools are not idempotent?
What if it generates a recursive loop?
Also, if I really want to do this, and if my tools are safe, why don't I just do a raw Open AI / Claude call and get deno subhosting [1] or E2B [2] to run it?
[1] https://deno.com/subhosting
[2] https://e2b.dev/
> Actually it doesn't, many standard interesting metrics will break because long-polling is not a standard request either.
As a person who works in a large company handling millions of websockets, I fundamentally disagree with discounting the observability challenges. WebSockets completely transform your observability stack - they require different logging patterns, new debugging approaches, different connection tracking, and change how you monitor system health at scale. Observability is far more than metrics, and handwaving away these architectural differences doesn't make the implementation easier.
I'm not commenting about the white boarding.
The article is about the fact that if you did just this, you're at a disadvantage because people use Chat GPT to cheat on these and get ahead.