HNHacker News
TopNewBestAskShowJobs

wittydeveloper

209 karma · joined July 12, 2012

submissionscomments
wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
Nope, you can also deploy and manage your own browsers as soon as they accept custom extensions.
wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
Does "agent-based" here mean a CLI/Coding agent or software-level agents you built? I'd say that Stagehand is subtly faster for local use cases (batching, in-extension design) while covering both CLI (via https://browse.sh) and framework-level use, whereas agent-browser is CLI-only.
wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
The cache serves as a managed server-side layer in the Stagehand API, keyed to instruction, page content, and the options you provide. Model configuration is intentionally excluded from the cache, so switching models does not invalidate it. If the page structure changes, the call does not result in a cache hit, and we revert to full inference. Additionally, any cached selector that no longer resolves will also revert, with self-healing disabled during replay.

Regarding the local file issue: currently, caching keys are based on a Browserbase session. This means that with a purely local browser, the cache option has no effect, and there’s nothing to verify. Each call returns metadata (including status, miss reason, threshold, count, and tokens saved) so you can observe churn rather than make guesses about it.

wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
We removed agent because so many great harnesses are available in the ecosystem. Instead of keeping it, we decided to make Stagehand v4 better integrated with popular harnesses, both at the Coding Agent level (Codex, Claude Code) and frameworks level (Eve, Deep Agents, Mastra, etc)
wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
Exactly, as Stagehand now runs inside the browser, you'll save on the round trip. Also, enabling batch actions will further accelerate your test suite.
wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
Nope, Stagehand is open-source and works with local browsers by default.
wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
A lot, especially for performance. That's why we built our own benchmarks that account for both the model and the harness. For example, with Claude Opus 5, the accuracy gap can be up to 3% and performance up to 200ms, depending on whether you're using Deep Agents, Eve, or Fx.

More details here: https://www.stagehand.dev/evals

wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
It still uses CDP but communicates from an extension within the browser instead of a script running in a separate runtime or, worse, in a separate region.
wittydeveloper··on We made Playwright 2x faster and 80% more token efficient
We built Stagehand 2 years ago (24k stars and 4M monthly npm downloads) and recently fixed its biggest flaw: round-trip latency.

Every action performed requires a round trip between your script and the browser (short when running locally but increased when running in the cloud). We also saw multiple posts complaining about the eager token appetite of Playwright MCP.

For this reason, we rebuilt Stagehand from the ground up and shipped v4, where Stagehand controls the browser from an extension automatically loaded upon your browser startup.

Stagehand v4 comes with batch command support, dedicated token-efficient methods `act()` and `extract()`, and a brand new architecture making it 2x faster than Playwright and 80% more token efficient.

You can see for yourself by looking at our benchmarks, comparing its performance across a dozen models (frontier and open weights) and tools (Codex, Claude Code, and more): https://www.stagehand.dev/evals

Ask me anything!

wittydeveloper··on Show HN: StepKit, an open and cross-platform durable execution standard
You can see durable execution as a combination of persistent state and queues (simplified example). With regular queues, the state is spread across many places from messages, runtime and external storages where the primary value is the reliability of the message processing and simple error management. Durable brings more advanced error management and end to end reliability with persistent state.
wittydeveloper··on Replicating Cursor's Agent Mode with E2B and AgentKit
Yup, I definitely faced multiple loops or "blocked" scenarios when iterating on the prompt and tools. The most obvious one is when the Agent generates a terminal command that requires some interactivity (ex: "confirm to ..."). This causes the Agent to build all files from scratch adding a 5x more steps than necessary.

I'm planning to try to make terminal STDIO work as a "human in the loop" pattern or even better as an "Agent in the loop" one.

wittydeveloper··on Replicating Cursor's Agent Mode with E2B and AgentKit
E2B does support Python, AgentKit is only available in TypeScript for now
wittydeveloper··on Replicating Cursor's Agent Mode with E2B and AgentKit
Not the version featured in the blog post, but adding a MCP server like Stagehand could enable it to search on the web!
wittydeveloper··on Replicating Cursor's Agent Mode with E2B and AgentKit
Claude Code is a complete tool that is ready to be used locally, focusing on Anthropic API features, while this coding agent is for educational purposes (like OpenAI's Swarm) and works with any LLM provider. Thought this Coding Agent could also be used in a real setting as a Pull Request or GitHub Coding Agent!
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
1. Yes, we currently provide a simple approach to retries but are already working on providing features for idempotency.

2. All sensitive data (tokens, env vars) are encrypted on our side, however, we don't prevent users to print their values in the logs yet - it's planned in our next releases.

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
It is actually comparable to both.

We aim to provide a similar feature set as Sidekiq (throttling, unique jobs) but with a complete hosted solution.

While Defer and Temporal can both be used for writing background jobs, workflows, and CRONs, there are some core design differences.

Temporal has been created as the Kubernetes of highly distributed systems, enabling developers to write code that runs on multiple regions without worrying about possible termination of the program and interruption of workflows spanning across multiple steps.

While Temporal can be used for background jobs, workflows, and CRONs, its main goal is to ensure that highly distributed tasks will reliably be executed. That's the main reason why Temporal API is so verbose, with many concepts to deal with.

Defer, on the other hand, provides comparable reliability while focusing on the developer experience.

You can write workflows, CRONs, and jobs that run for hours without worrying about them being terminated.

All this, with a simple API that enables you to write some workflows (background functions calling other background functions) in plain TypeScript, with no mental model to fit in.

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
You're right! We will update them soon.
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
Thank you for your feedback! Can't wait to onboard you on Defer
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
Thank you!
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
Thank you! Join our community so we can help you onboard and get updates on the upcoming features (more is coming soon): https://discord.gg/x2v84Vqsk6
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
When deployed in non-local environments, background functions get executed on our infrastructure with their (securely stored) arguments, spawn in a dedicated isolated temporary container.

For our users concerned about data locality, we recommend pushing the minimal data as arguments (ex: ids or external ids) and fetching the data during the execution on our side, if needed, through a dedicated SSH tunneling setup. Once an execution is done, its associated isolated container - gets a dedicated VPC and disk - gets destroyed permanently. We will also provide on-premises solutions for Enterprises.

In local environment (dev), background functions run completely synchronously and locally, no call is made to Defer.

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
We couldn't agree more!
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
We support both CJS and ESM but showcase examples in ESM in our docs and landing
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
Here’s how our Builder works: When a commit is pushed, we fetch your application’s repository from GitHub and compile (if TypeScript) all the files in the first `defer/` folder found. Then, we require each of those files to retrieve the metadata exposed by the `defer()` helper, on the `default` export (is it a CRON function or not, a function name, concurrency, retries, etc).

When a background function gets a call from your application, the `defer()` wrapper intercepts this call and pushes an execution to the Defer API with the function’s name and serialized arguments.

I hope it makes thinks clearer, let me know!

[update: grammar]

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
I agree with you, as mentioned in the article, I’ve built similar solutions as yours at Algolia, using Redis and Kubernetes. However, not all developers know and want to do Kubernetes or RabbitMQ and not all companies can afford to invest time to set up and manage it (https://docs.bullmq.io/guide/going-to-production).

We address privacy by encrypting all the data on our side (doing a second pass with a symmetric PGP key for tokens such as GH tokens, and environment variables) and advise companies that want to keep their data on their infra to push as minimum data in arguments while leveraging a dedicated SSH tunneling setup between our infra and theirs.

When it comes to the SLA/up-time of home grown, my POV would be that, again, achieving good results on those often requires SRE engineers, which is an investment.

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
Sure!

We got our first customers from both channels: network (sales) and inbound (twitter discussions, etc). I agree that top-down is not working well for a newcomers but getting better at a later stage.

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
Sure, please note that was referring to “modern development standards” in the Node.js ecosystem.

Node.js, with the flexibility of the JavaScript language, allowed the rise of great abstraction and other domain-oriented API designs, a bit like Ruby on Rails did with Ruby.

The arrival of React.js Server Components pattern enabled the isomorphic pattern (first applied to mobile and front-end apps) to reach the server side of things with patterns such as Server side rendering, popularized by frameworks like Next.js or Remix.

Those new patterns and abstractions allow Node.js developers to move faster while building more complex applications to match users’ requirements: real-time, performant apps, richly integrated with third-party products.

Beyond code, those new coding habits came with new products such as Vercel or Supabase that help them to get the infrastructure done in no time, without any DevOps knowledge (good article on this topic: https://vercel.com/blog/framework-defined-infrastructure).

“modern development standards”, applied to Node.js, do not only apply to coding experience and productivity (ex: the rise of monorepos, TypeScript, SSR) but also to enabling developers to configure their infrastructure from the code.

[update: typo]

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
The first blog post is actually pinned but an icon or label to highlight this is missing, you're right! I'll fix it soon.
wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
We primarily focus on helping Node.js developers to build products without investing time in infrastructure while keeping it configurable.

As @thdxr mentioned, our goal is also to provide top-notch queueing/scheduling features that are complicated to achieve with Lambda/SQS/Cloud Functions, such as dead letter queue support, throttling, or a product-oriented dashboard. In the same idea, we plan to provide better integration, for example, as you mentioned on secrets management by integrating with Doppler, AWS Secrets Manager, and more - the same goes with linked deployment pipelines.

wittydeveloper··on Launch HN: Defer (YC W23) – Zero-infrastructure background jobs for Node.js
Yeah, that will be the way to go. This is actually how on of customers is achieving this behaviour for polling analytics.

Right now our Executions list is not ideal for such pattern but we will soon release filtering based on arguments which will help to get all the executions linked to a specific sequence.

Page 1 of 2Next →