Kitesurf: Agent-first browser that runs in V8 isolates
blog.cloudflare.com
blog.cloudflare.com
(I wasn't involved in building kitesurf, but I am informed that they intend to open source and upstream their patches)
[edit: for others reading who don't usually nerd out on browser automation protocols: webdriver bidi is the new-ish w3c cross-browser standard inspired by CDP - the main magic was the upgrade to websockets and also to standardize the capture of network-level traffic. there are still feature gaps between CDP and BiDi (in spec and implementation), but long term, i believe we should bet on web standards, not proprietary protocols controlled by one company.
(disclosure: i started the selenium and appium projects.)]
(if kitesurf does upstream their patches then presumably we'll get a CDP-based automation API as part of that)
Totally agree.
Not sure if you're involved in the development / spec process for WebDriver Bidi, but the big limitation atm is that it has almost no support for the devtool inspection use cases served by the Chrome Devtools Protocol (CDP) and the Firefox Devtools Protocol (FDP).
The Servo and Ladybird browsers both have FDP implementations (and Blitz has an in-progress CDP implementation) for this reason. But we'd all love to switch to a single standardised protocol if it had the requisite support.
Just curious on your thoughts about how Webkit was architected then, I guess it's not a modular system where you can separate out things like "Localstorage" support?
Yes, it's new engine separate to Webkit/Blink/Gecko/Servo/Ladybird/etc
> Just curious on your thoughts about how Webkit was architected then, I guess it's not a modular system where you can separate out things like "Localstorage" support?
Honestly, I'm not super-familiar with Webkit's architecture. It's a huge codebase, and it's also C++ which is always pretty intimidating. I believe Webkit is more modular than most of the others, but Blitz goes quite extreme into modularity:
- The core is not coupled to the HTML parser
- The core is not coupled to the networking
- The core is not coupled to the rendering backend
- The core is not coupled to the windowing/input layer
- The core is not coupled to the JS/scripting engine
- The style engine (Stylo - shared with Servo and Firefox) is mostly implemented as a library which can be used independently
- The layout engine is mostly implemented in two libraries which can be used independently of the rest of the engine (Taffy for Flexbox/Grid/Block layout and Parley for Text/Inline layout)
So, yes I'd hope that it will be possible to individually opt-in to features like localstorage (once we implement them), but it goes a bit further than that.
Is it a good idea to already build something on top of Blitz?
These two feel like they are opposing teams, I don't think they are colluding today, but how long will that last, this seems very suspicious I say that as a long time cloudflare user, I welcome making the platform agent friendly and adding agent specific deployment cloud stuff like Cloudflare OS is something I can live with as well.
But this is going a bit too far, what's next AI bot net to scrape content from sites protected by Cloudflare? I don't want to sound entitled but man do we deserve better.
They're cutting AI off at the legs for everyone else, then building the new tool to sell AI enablement.
They'll probably let their customers pay to bypass their protection scheme, which will complete the loop.
It really rubs me the wrong way.
It should 100% be a different company. I don't feel safe with Cloudflare being the ones building this tech.
> Run headless Chrome on Cloudflare's global network for browser automation, web scraping, testing, and content generation.
Does Cloudflare the CDN allow these browser instances to bypass their own anti-bot mechanisms? Or will Cloudflare the CDN block them the same as if someone was running scraping bots from a different provider?
Will Kitesurf in Cloudflare workers get special bypass privileges to content protected by Cloudflare the CDN?
We also have a documented UA and sign our requests with Web Bot Auth: https://developers.cloudflare.com/browser-run/reference/auto...
My wife really dislikes building up the shopping cart for our weekly grocery delivery, so I built an agent... thing with earendil's npm libs. It takes the menu my wife has decided on, confers with her about the ingredients (if it hasn't seen a recipe before), and then uses Chrome's devtools protocol to head to Walmart and add everything to the shopping cart.
It works fairly well and uses the local models I have running on my Mac Studio.
That is fantastic. Last I checked models capable of running on commodity (anything below a dedicated GPU rack) hardware were very lackluster.
I could probably drop the smaller Qwen at this point, but when I was first building this I was having an issue with search results and cart data filling up the main agent's context.
There's many parameters in calculating a risk profile including having a residential IP number. I think you have that.
I didn’t “use an agent to find a receipt” in the sense that I purpose built one. I just asked my existing agent that I talk to on telegram by photographing the thing I wanted to know if we could return and while I changed the baby it chugged along and by the time we were ready to go it could tell me whether we did buy it at Costco and when so I know if I can return it.
tldr; to workaround the lack (or shortcomings) of public m2m APIs in web apps
- Apple Appstore Connect (gazillions of forms of metadata to release an app) - AWS - DigitalOcean - Google Play Store
Whenever I dread logging in because I know the simple sounding task requires me to click through countless menus I use an agent browser. With confirmations of course. However, while the agent clicks through these (oftentimes dog slow) UIs I can do other things. Once it requires permission, I read, decide and act.
I also tell my agents to remove annoyances from websites I browse, rearrange the content so that it's easier for me to view. For example when somebody publishes a table where they compare their newly released AI model to others I tell my agent to highlight highest result for each benchmark in every table on the page. I could do it myself with a bit of JS but why bother if agent can write it for me. I added a functionality to my agentic browser that lets the agent make userscripts for me that I can trigger with a push of a button.
I also ask agents whether the specific information is on the page that I'm currently browsing in language I don't understand (or just among the clutter).
Once I asked agent to put more than a dozen items into a cart for me (which names I pasted) because the ecommerce site didn't have convenient way of doing that.
So basically Grease Monkey on steroids + TD;DR;whaat?
Local Qwen3.6 is smart to do all that but I have option to switch to remote stronger models.
From what I understand, Kitesurf eventually went in its own direction and isn’t simply Obscura running on Workers. Still, knowing that Obscura helped get the original experiment started means a lot.
I recently added native rendering to Obscura. It can now take screenshots, stream screencasts, and generate PDFs without Chromium. The repository has also passed 21,000 stars.
For context, I’m 16 and have mostly been building this with my friend, so I’m still figuring out what the project should become and how to keep developing it sustainably.
Happy to answer any questions about Obscura or the rendering work.
time for another approach to run your agent's web searches through this (or by mocking browser signature), with potential cf bypass built-in!
We have been keeping a close eye on BiDi as well.
It's a web data tool, but as something not used for browsing, by definition this is not a browser.
a welcome addition although it'd be very easy for websites to fingerprint and block
> One last thing: we're going to open source Kitesurf once we're ready — hopefully soon. Our goal is to let any customer deploy their own version of Kitesurf on their own accounts, if they want to.
You can block today, Kitesurf doesn't try to hide.
https://developers.cloudflare.com/browser-run/reference/auto... https://developers.cloudflare.com/browser-run/faq/
Results from the last hour, verbatim:
- lemmy.world /api/v3/site: registration_mode RequireApplication, captcha_enabled true, require_email_verification true, and an application question that explicitly rejects temporary email. Three independent walls on one signup. - lemmy.today and lemy.lol /api/v3/user/register: {"error":"captcha_incorrect"}. The captcha ships as base64 PNG plus WAV, so it is a wall for anything without a decoder, headless browser or not. - bsky.social com.atproto.server.createAccount: {"error":"InvalidPhoneVerification"}. - Publishing, by contrast: api.telegra.ph and write.as both take an unauthenticated POST and hand back a public URL.
A browser in a V8 isolate does not help with any of the failures above, because the gate is a CAPTCHA, an SMS, or a card on file, and an isolate has none of those. The same is true on the payments side: an agent can hold an address and receive, but every write path in that ecosystem is a signature over a payload, so if something else custodies your key the machine-payments world is read-only to you.
The missing primitive for agents is not a browser. It is a portable identity and a spendable balance that are not borrowed from a human's phone and credit card.
Full map of what was reachable and what was not: https://write.as/ih3l0kd78lpb1
It's funny you should mention that since Cloudflare is also working on bot identity and payment (x402).