Prove you are a robot: CAPTCHAs for agents
browser-use.com
browser-use.com
Definitely chinese.
In Japanese, they say 'hundred' instead of 'one hundred' "百 二十 一"
"Kanji", 漢字, in Japanese literally means "Chinese character".
The kana, hiragana or katakana, are only used in Japanese writing.
The people behind the website asked a voice agent to program it, and the STT parsed "agent" as "asian."
API Error: Claude Code is unable to respond to this request, which appears
to violate our Usage Policy (https://www.anthropic.com/legal/aup). Please
double press esc to edit your last message or start a new session for
Claude Code to assist with a different task. If you are seeing this refusal
epeatedly, try running /model claude-sonnet-4-20250514 to switch models.If a few tool-wielding humans slip through, that's fine (traditional CAPTCHAs also let in our stealth agents)
Humans can use agents behind the scenes to crack it, right?
My best guess is that this a way of making a system talk to your agent without you knowing what they are talking about ? As a way of not exposing the real sign up method ?
@powforge/captcha (MIT, npm) has two layers: SHA-256 PoW for anonymous visitors (~5 seconds), and an L402 Lightning payment tier for agents that want to identify themselves. An agent that holds a Lightning wallet and pays 3 sats is providing a different signal than one that grinds a puzzle — it has economic stake in the interaction.
The Lightning tier hasn't seen any payment yet in our test deployment, but the architecture is live. The interesting question going forward is whether agent-to-agent surfaces will converge on capability proofs (something like L402 + identity scoring) vs puzzle solutions.
The second is that if I hit L on Chrome for Mac OS on the linked page it takes me to their signup page (presumably because I have no account). So that's a keyboard shortcut to take you to the browser-use app page. But why 'L'? And it's funny that Cmd-L (focus address bar and select address) in Chrome triggers the L effect but does not in Safari (where L on its own still works).
Initiation Mathématique (1906) by Charles-Ange Laisant (1841--1920), number 53. Le chien et les deux voyageurs.
The setup here has two pedestrians walking in the same direction with a dog running back and forth between them. One of them starts out some distance ahead of the other but, because the one behind walks faster, they eventually intersect. It briefly mentions a variation where they are walking toward one another, as in the typical trains & fly version of the problem. Best of luck finding older, I wouldn't be surprised if it's out there!
It's also infamously the subject of a Von Neumann joke
>Two bicyclists start twenty miles apart and head toward each other, each going at a steady rate of 10 m.p.h. At the same time a fly that travels at a steady 15 m.p.h. starts from the front wheel of the southbound bicycle and flies to the front wheel of the northbound one, then turns around and flies to the front wheel of the southbound one again, and continues in this manner till he is crushed between the two front wheels. Question: what total distance did the fly cover ? The slow way to find the answer is to calculate what distance the fly covers on the first, northbound, leg of the trip, then on the second, southbound, leg, then on the third, etc., etc., and, finally, to sum the infinite series so obtained. The quick way is to observe that the bicycles meet exactly one hour after their start, so that the fly had just an hour for his travels; the answer must therefore be 15 miles. When the question was put to von Neumann, he solved it in an instant, and thereby disappointed the questioner: "Oh, you must have heard the trick before!" "What trick?" asked von Neumann; "all I did was sum the infinite series
curl -fsSL https://browser-use.com/cli/install.sh | bashApplication error: a server-side exception has occurred while loading cloud.browser-use.com
Great first impression!
Is this just a marketing stunt?
Silly solutions for silly problems :^).
So, showing true agent to agent interactions is interesting, but one could never be sure that's what you were actually seeing unless you were in control of all the agents.
If a human uses the API key after, that's fine. You also get access to our free tier if you sign up the traditional way clicking around in the UI
Holy shit - why don’t they produce an AI summary and plonk it in there for everyone to use? The energy savings across all people who’ll read the summary would be staggering!
Which LLMs best drive these? Claude/Gemini, etc., or is anything local actually competent at it?
Can they understand layout and visual cues with a VLM or multimodality?
Are they robust enough to interact with threejs and videos and whatnot, or can they just blindly navigate the DOM?