20 karma · joined July 22, 2020
We are closer to infrastructure than a library or framework; we give developers a live agent they can chat with in a single API call.
1. Our approach has cron-based or trigger-based automations built-in. Building automations with claude agent sdk requires setting up separate infrastructure.
2. Our approach has self-learning built-in. Building a feature like "dreaming" https://docs.openclaw.ai/concepts/dreaming with claude agent sdk also requires setting up separate infrastructure.
3. Our approach decouples the harness and the compute, which lets developers enforce a stricter security boundary, while claude agent sdk ships with the harness, shell, and filesystem in one process https://platform.claude.com/cookbook/claude-agent-sdk-07-hos....
4. Our approach does not vendor lock developers.
You could pick the latest harness and then switch when another better one rolls out. Our bet is that a developer's time is better spent speaking to their customers than switching harnesses.
If a startup has a specific flow they want the agent to take and their traffic is bursty, then I'd recommend using a framework like Mastra and deploying onto a sandbox.
For long-running always on agents where it's important to learn the users preferences overtime, our approach is the highest ROI.
Cost is the token usage and container uptime.
> One Docker container per-customer sounds like it would be really expensive.
The advantage is per-user memory and self-learning. For context, Claude Managed Agents uses one sandbox per session: https://platform.claude.com/docs/en/managed-agents/environme....
> Are they started on-demand, or run 24/7?
24/7 (best for customer-facing chat products).
> What keeps users from using the agents for general purpose tasks, protects against prompt-injection, etc?
Users define their agent with a system prompt, tool definitions, and skills (which separate a media generation agent from a people search agent). We use Openrouter which has a prompt injection detection feature: https://openrouter.ai/docs/guides/features/guardrails/prompt....
Even as the cost of writing code goes to zero, those two pieces of information are non-commodities.
By providing Hermes with a system prompt, custom tools, and skills, developers get the agent loop, session management, automations, sandboxed deployment, and self-learning for free.
You can give Prism a try at https://prismvideos.com - I'm excited to hear your feedback.
Nevertheless, things are trending more in this direction, and AI influencers will soon become the norm. Brands should be required to disclose when their marketing is AI.
It's worth mentioning that AI videos on Prism (and on any platform) do not have to be purely prompt to creative. For example, a brand designer can take an existing creative for a billboard for example and then use AI to generate images of this creative at a train station, in the Louvre, at a bus stop etc (without actually going there and shooting images).
We've found that Seedance is good at photorealitic faces, Kling is fantastic at generating audio (highest quality model in terms of syncing character's face to the words they say imo), and Sora is great at UGC.
Would be curious to see your script.
One risk with these new standards for agent auth - which we will of course support if our customers want it - is that the websites that need them the most are the least likely to adopt them.
The main use cases for browser agents are for paying utility bills on old government websites or finding receipts for an expense report on a website without an API. There is a no reason to use browser agents on a website like Linear for example. A developer is better off integrating via API or MCP.
Therein lies the main challenge; the websites where browser agents are most useful are the same websites that are least likely to adopt new technology (it was their not adopting new technologies that made them good candidates for this browser agents in the first place).
I think this new standard is awesome, but I fear that the websites that support it will be those websites that didn't need it in the first place (because they could just as easily add an API).
When our agent signs in, we input the forwarded otp code to get access.
Right now, we're focused on building connectors for our customers, which has not yet involved Captcha solving.