Meta's Muse is fantastic for web scraping
sigh.dev
sigh.dev
Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.
I agree with you, though, that the future has almost certainly already been decided, and see the eventual outcome of all this selfishness to be mandatory device attestation to access most services on the internet—which will of course never be allowed on any open platform. AI engineers who believed in freedom to compute (or just individual freedom more generally) have committed perhaps the greatest self-own in the history of our species to date.
Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.
Humans are interesting creatures.
Edit: Please before assuming that I'm assuming things, this is an observation I'm making over time. It's possible that I'm in a bubble, but it's not a sample size of 1 (i.e. The comment I replied only).
Throttling is poor help though. Mass scrapers are using "residential proxy" loophole + rotating UA and other attributes. You can't throttle somebody without identifying them. Unless you're talking about a global rate-limit.
Once the LLMs create sockpuppets to get around that, the web services will need to resort to profiling users more aggressively so that they know which actual human an account corresponds to.
If someone has a malicious browser extension that uses their session to scrape Reddit then, they're probably going to see significant usage obstacles.
We are headed to a very user-hostile place.
It's not that "we" are going from one thing to another. It's that these are two different people, with different ethical boundaries.
Is defeating captchas unethical? I don’t think so. Not on its own.
Is scraping unethical? I don’t think so. Not on its own.
Are there tons of uses for both of those that are sketchy or outright wrong? Yeah, absolutely.
As an example, imagine a really small community maintaining a small site/wiki/cms/forum that has the ultimate goal of promoting human relationships around a common interest (let's say retro computing as an example, but it could be anything). There are many many many such communities on the internet.
It's very improbable you will specifically instruct your LLM to access their site, it's way more probable it will happen without you even knowing, as a result of you doing some /deep-research or something. And not only that, but your LLM will probably spawn a ton of agents to gather as much information as possible in as little time as possible. A torrent of requests will go at this community's site, effectively killing it. They are a small community, they use their spare time and money to maintain something to serve them, they don't have the resources to serve your LLM and until you showed up they probably never even had to think about Cloudflare. They are certainly not against you getting the content, but they don't want you causing them issues either and you just did.
End result? Your LLM (effectively you) DoSed a small community's site. You caused harm. Could you have caused similar harm if you were doing it on your own? Sure. But you would have done it on purpose, not accidentally while instructing your LLM to do something else.
This isn't a made up story, it has happened already more than 1 times.
So the question is, now that you know your LLM can cause harm without you even knowing it, how does this change your stance?
We can stop putting information in the open in any form, as well.
FAIR: Findable, Accessible, Interoperable, Reusable.
Considering that, shall we paint the boards black or lock the doors to community spaces?
Also, while I assume this is not your first account here, these are worthy of reminding:
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Throwaway accounts are ok for sensitive information, but please don't create accounts routinely. HN is a community—users should have an identity that others can relate to.
For more, please refer to https://news.ycombinator.com/newsguidelines.html
Sure, Nazi, Hitler or Stalin all considered themselves to be the good guys. Kushner, Trump, etc are also acting within their ethical system (if I am getting money or power or fame it is ok to do it). Just about the only exception is Thiel who openly frames himself as evil, but is proud of it.
That does not mean we cant criticize their crappy actions or "ethical frameworks".
And yes, all the above are making deliberate decisions to be unethical assholes.
the only solution is to drive your regular browser with all your sessions/cookies via an extension
This is exactly what I'm doing, but mocking/randomizing all the sessions/cookies/params (like resolution, OS, WebGL , etc.) in a separate Chromium binary. It's popular these days, but imo using your normal browser for agentic stuff is a very bad idea. These models do dumb stuff all the time.
For testing proxies should be used to avoid IP ban issues but given enough time modern agents can figure out how to bypass most of the modern scraping/automation prevention mechanisms.
I have built and used a lot of different automations and web scraping implementations for my business and it's never got permanently stuck yet, some take a bit longer, some shorter, but all within a reasonable time with little external help they have succeeded in their tasks.
Is that necessary? You could start Chromium with an open debug port and use Chrome Devtools Protocol to send commands.
I've had success launching the browser and using dumb dumb methods to get around the captcha before attaching playwright.
Dumb dumb methods = wmctrl, xdotool, bash (work fine for Cloudflare's "are you human" check)
Assuming "more people catching onto" this, expect Cloudflare and most everyone else to follow the suite.
> not sure what you mean by "Chromium"
I meant that literally. This thing: https://www.chromium.org/getting-involved/download-chromium/...
It searched for "Aliexpress" on Google and clicked the link, hence the Google UTM params.
Some days I'm literally fatigued enough I can't clear the overly complex hCaptchas these days I keep getting forced upon me and bots sail right through this crap.
But I'm wondering: at what cost?
(×) and their AI delegates
So of course I clicked yes and it dutifully convinced the site that it was not a bot.
Muse's utility has significantly dropped with the blockages.
To become truly useful again they will need to use residential proxies, but I can't see them use those due to the risks and reputational damage.
Spamhaus is nearly thirty years old and the notion of electronic distribution of IP and domain "reputation" lists is at least that old.
I'll bet my hat that the Internet "advertising" [0] industry has been calculating and determining the reputation of individual households (if not individual users) for at least a decade.
[0] The scare quotes are because its primary purpose these days is for dragnet private-sector surveillance.
One thing with Muse though is that you can scrape Meta's own sites which are currently all going under login walls and restricting discovery/search entirely.
Multiply it by 1,000,000 access attempts - daily.
It is question of time when even normal browsers would need some allowed fingerprint to go to sites otherwise caught by clouflare or similar wall.