Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped
techdirt.com
techdirt.com
Google does supply search grounding with Gemini API calls, and that is handy, but not general enough.
https://abcnews.com/Technology/wireStory/eu-forces-google-sh...
> Judge Mehta said in the 223-page ruling that Google must share some of its search data with “qualified competitors” to resolve its monopoly.
https://www.nytimes.com/2025/09/02/technology/google-search-...
It's up to sites if they want to require an open standard and risk losing clients that haven't implemented it yet.
>churning constantly to anticompetitively maintain their monopoly.
Google open sources the implementations of these. Competing browsers like Brave and Edge are able to integrate this open source code to support them without themselves having to deal with constantly implementing new features.
Why don't you make a de-JSing Google proxy? Like SearX-NG?
Meanwhile Google actively blocks and attempts to pursue legal action against scrapers. Not to mention prohibiting it through TOS (getting your Google account revoked can be life-ruining for many people).
I stopped using Google Web search as well after this point
NB. The autocomplete endpoint at clients1.google.com for example does not require Javascript, nor HTTPS
One could take the Google suggestions and search those strings in other search engines. Would the search results differ
Google Web search might be useful for finding "popular" results as these are the only results Google LLC management wants to show its ad targets, so-called "users", because funneling ad targets to the same sites ultimately creates larger audiences for advertising and more spend from advertisers
But if one is searching for "unpopular" results, e.g., performing "discovery", then using Google Web search is, IME, certainly not the best method of searching
IME, quitting Google leads to more creative search strategies; I have found stuff that I never would have discovered using Google Web search
NB. Google Web search requires Javascript. Scholar search, News search, etc. do not
An HN favourite:
I’ve stopped using Google search as well but JavaScript is the least of their sins.
It’s completely reasonable for a website to require/leverage a technology that literally 100% of their customers have.
"Allowing javascript can create security vulnerabilities (e.g. in the case of infected ads), and it is used for browser fingerprinting, so that users can be tracked and ad targeted across sites (especially if the information is sold, as part of a database, to third parties who combine it with other databases) even when they deny cookie and local storage permissions and aren't logged in to a google account and are using a vpn or proxy. If you go to one of those "how unique is my browser" test websites, if you have javascript blocked, they have comparatively almost nothing to go off.
It also always loads slower and uses more data and ram than plain HTML, and unfortunately many people in rural areas still have shitty internet service and/or old hardware. And it's rubbish if you just want to browse the internet on something like an ereader or an old laptop.
Unless you specifically need features that aren't possible without it, it's a worse experience, and Google is choosing to make it mandatory for the whole site instead of just for the specific features that need it because it's good for their ad revenue. It's a means for Google to track and sell the data of users who are trying to opt out and preserve their privacy.
I haven't found a single other search engine that does this. I'm pretty sure even Bing doesn't (although they do their "log in for reward points" thing instead)."
Why would a user prefer in some cases not to use Javascript from an adtech company
Depends on the user. For example, here's one user who publishes a website on the subject
https://disable-javascript.org/
Consider the level of data collection and behavioral surveillance enabled by the SearchGuard Javascript at issue in Google LLC v SerpApi LLC, as described here:
https://searchengineland.com/inside-google-searchguard-46767...
Google LLC is an adtech company
Google LLC may claim that this Javascript is an access control to protect material copyrighted by others
But the court in Google v SerpApi has found that to be false, with the possible exception of "Knowledge Panel" material. The court found that Google LLC is not authorized to protect the returned URLs in a SERP as copyrighted material
If the purpose of this Javascript is not access control for copyright protection then what is its purpose
Can the user control this Javascript. No, it's under the control of Google LLC
Can the user control how the data collected by GoogleLLC via SearchGuard or other Javascripts is used or where it may be sent. No
Mitigation may require disabling Javascript
https://captaincompliance.com/education/tracking-technologie...
https://medium.com/@Kevin_Finnerty_Gabagool/adtech-the-silen...
https://arxiv.org/html/2508.07454v2
https://adguard.com/en/blog/weblock-location-tracking-survei...
Many of these tactics, including Real-Time Bidding (see AdGguard blog post), as implemented by adtech companies such as Google, require Javascript
It is reasonable that a user might prefer to avoid running Javascript from adtech companies, such as Google LLC
Some of us have been searching the www since before Javascript existed
Using others' Javascript is a choice the user gets to make. Adtech commpany Javascript for web search should be optional. Costs may outweigh benefits. Depends on the user
*company
Edit: now I'm worried. Is JS actually disabled?
Typical behavior from a big company with immense resources. They probably thought they would get a settlement or SerpAPI could not afford to fight. I assume they are pretty small, at least in comparison to Google (I've never heard of them).
Google has so much money that even a "loser pays" requirement on litigation probably would not disuade them.
In USA copyright requires a minimum degree of original creativity in the selection, coordination, or arrangement of the data.
I think it's a rather grey line to say that Google search results are just facts, but eg maps are copyrightable. There's a rather large amount of effort involved in crawling and ranking the web - the PageRank itself should be copyrightable.
I don't think a map is a good example. A picture is probably a better one. A map, by its very definition, is not a replica of any part of the original artifact. Not merely because a map is not the territory, but also because it's not even a direct, unaltered view of the thing. There is clearly some creativity required in putting together a map, since it requires you to decide what to include, what to leave out, what to exaggerate, what to distort, etc... as evidenced by maps of the same area looking vastly different.
By contrast, search engine results are stitchings of various pieces of the text on the page, verbatim. How much creativity that embidies is probably akin to that within a photo.
[0] - I believe a specific case was Nintendo vs Prima publishing, which was even involving a map of a fictitious construct.
Yes I remember reading a few years ago about the Royal Australian Navy trying to find Sand Island which I think was some tiny island marked on maps of part of the Pacific Ocean off the coast of Australia - what they found was open ocean 1,100m deep, and concluded it was the map-maker's "mark".
An example is courtcase against Meta for using a Australians mining billionares likeness to promote a crypo investment scam.
https://www.afr.com/technology/dad-it-s-a-fraud-call-that-sp...
I say those words in quotes - they have no intention of shipping you anything, just skimming low-hanging fruit from someone's ideas.
I'm kind of surprised that there haven't been criminal cases.
As soon as a dollar of profit is involved in a content moderation decision, a business should be fully liable for the decisions around content on their platform. If I report an ad to Facebook and Facebook decides to keep it they should be accepting legal responsibility for that ad.
You want to stop scams online, you make the platforms liable and then grant them the ability to recover the losses by going after the advertisers.
1. Takes money from the scammer
2. Tells the scammer how to target people
3. Serves the scam to the consumer, ensuring they see it
4. Takes a cut of the action when the scam is successful
5. Tells the scammer how to optimize their campaign
6. Continues working with scammers after they're reported
The scammer has a fairly small part in the overall operation. The platform is doing almost all of the actual work perpetrating the scam. They're not remotely innocent here.
Your whole chain stops being relevant if you actually persecute criminals with the same gusto as Disney jackboots anyone voilating their IP.
Instead you demand megacorps to start scanning content and censoring people.
For what it's worth, I wasn't demanding anything. I was solely pointing out that the platforms aren't neutral here. They're active participants in perpetuating these scams.
You are also free to call it censorship, but I am not aware of any jurisdiction where fraud is considered protected speech, so as far as I'm concerned, censor away baby.
If they choose to get out of the business of facilitating crime because of that then so be it.
When Google has no profile on you, your view is virtually worthless, so it's only bottom feeders that bid on those views.
Average users get Coke and Tide ads. Its usually the most technically adept that get the worst ads, and usually they just turn their ad block back on.
But when they turn off ad-block, to "see what it's like", they are not getting an accurate view. It's usually all crypto and other scammy/vices type ads.
Go look at your mom's browser. Her Google ads are going to be clothes, tissues, and cookware.
Nigerian prince scams cannot outbid Kleenex without breaking the economics of the scam. But they can get loaded when Kleenex no-bids because they don't know who the viewer is.
Trust me, ad-tech is far far beyond 2004 when ad block showed up.
(I'll add that of course there are other less legitimate ad networks, and I'm not counting "snake oil" products, which are essentially scams but customers still swear by them)
I don't understand why you claim that targeting doesn't happen; I see it happening with people I know, who wouldn't even know what adblock is.
But after that you would get dialed in and stop seeing them.
Google's core mission is to figure out what you are going to buy before you buy it so they can have you click through them to make the purchase.
They generally have negative interest in serving scams/malware ads because people generally don't want to buy those things. On the same token though it's a very hard problem to 100% solve, and the people impacted are usually the lowest value users anyway.
Ah, good old times.
Google respects if one doesnt want to get indexed by the crawler:
https://developers.google.com/search/docs/crawling-indexing/...
Basically, search engines are publicly scrapable, though I do wonder as from a law point of view, that it must be within the murky waters as to what a search engine means in terms of seperating its search engine code/its recomendation engine and the public data much of which are intertwined with each other.
I believe that the argument that could be made is that the recommendation engine is the way it is because of all the data and its unseperable to really copyright the whole mechanism in all its glory.
Speaking of which, it seems that AI models feel really similar. Does this judge lawsuit show that AI model weights aren't copyrightable as well? If a search engine is built on public indexes then so are the AI models. I was just writing similar comment on another thread but it seems to be the case, definitely worth a blog article or thinking more about perhaps this judgement by this judge itself in general as well, I just have a vibe that this judgement has pretty far reaching consequences in its impact.
"Accordingly, the Court cannot reasonably infer that such licensing agreements grant Google authorization to deploy technological measures to control access to the copyrighted content at issue."
"Under Iqbal's plausibility standard, Google must, but has not, alleged factual matter that raises the inference that copyright owners authorized the implementation of SearchGuard as a technological measure to control access to copyrighted content. Allegations that Google has licenses to display copyrighted content in the Knowledge Panel are not sufficient, without more, to raise that inference."
Is the "Knowledge Panel" (KP) in SERPs comprised of copyrighted content
If yes, do KP licensing agreements grant Google authorization to deploy technological measures to control access to the copyrighted content
https://www.stackmatix.com/blog/how-to-get-google-knowledge-...
Alphabet is an Anthropic investor
Looking forward to the Amended Complaint by August 10
https://searchengineland.com/inside-google-searchguard-46767...
It behooves them to create systems that gets content creators paid as without content AI can not stay relevant. It also behooves entrepreneurs and technologists to create systems that solves this issue.
- Google search is on the way out. I don't know any of my peers who use it anymore.
- Coding models make doing extreme depth of work possible.
- Hoards of unemployed engineers now have access to 10x coding utilities and are looking for things to do. They will start clawing away at Google products.
- Just the other day, someone cloned Google Gsuite and it looked awesome
- Drive and Search will also be fungible products
- I'm itching at the chance to build my own phone operating system, and there must be thousands of others who want to do the same.
- Chrome can probably be replaced (Firefox gained a whole percentage point last month)
I don't think Google is safe anymore.
Two caveats that I'll give them:
- YouTube still has network effects and probably can't be dislodged
- Google cloud isn't going anywhere
What do they use?
Of all my tech friends, colleagues, I'm the only one who uses Kagi. Another person uses searx. Everyone else uses Google.
Of my non-tech friends and colleagues, everyone uses Google.
Me, almost, too.
I have some friends using duckduckgo - not many and only until Google is the default again...
Even assuming no hallucinations, an LLM can replace some use cases for search but not all.
It did use Google to provide me with the answers though, sooo...
Imagine you’re in the car or hands free or disabled and you just want the question answered.
Such is the case for people trying to poison--ahem "influence" LLM results for "what is the best restaurant in $my_locale". You're playing a dangerous game.
Sure. But that's not what's happening here, yeah? Both examples are providing sources. Actually I'd say that the LLM is doing a better job at that. You click a button and see 9 URLs to 9 news sites.
I think for the "influence LLM results" fears, it would be ridiculous to argue about which multi-billion dollar American company can be trusted more not to manipulate you.
Does it show relevant snippets for each of the 9 URLs? If not, then it's as bad as Google's AI summary. It's not that rare that the reference disagrees with the LLM.
The value of the non-LLM search engine is that in the search results, you see the relevant snippet, and if you want more information, you know quickly which links to click.
Not saying there's no place for the LLM for many search engine use cases, but a proper search engine replacement it is not.
Sounds like you're saying that the LLM isn't any better than Google's LLM-created snippets. The claim was that Google has no moat against LLMs, and your evidence against this claim is Google's LLM generated content.
So what I'm saying is:
Google Search Results > LLM results
And yes, that includes:
Google Search Results > Google Summary on search results page
Google's UX is terrible - but a search engine (a tool that can put you right into contact with the FAQ or forum where actual experts are discussing your problem) is always going to be powerful.
If you want to know what the relevant pages are, you need a search index.
I can assure you that it is not. If you download any model off of huggingface, it does not also include an index of the internet.
That screenshot shows the model making a tool call to an external search index.
Prompt: Who won the 2026 World Cup?
Answer: Spain won the 2026 World Cup, beating Argentina 1-0 after extra time in the final at MetLife Stadium on July 19. Ferran Torres scored the only goal in the 106th minute, coming on as a substitute in the 62nd minute. It’s Spain’s second World Cup title, having also won in 2010.
Prompt: Nearby BBQ places open now?
Answer: (a geolocation permission request prompt for the browser followed by) Right in [redacted] both [redacted] (4.6 stars, open until [redacted]) and [redacted] ([redacted]) are close and currently open.
A bit further out but highly rated: [redacted]
Seems like LLM does a good job on those questions…
The only issue is that the eyeballs stayed with the LLM, which allows it to hold a position that is potentially very powerful. This is particularly problematic with people who believe that computers are infallible.
But the user doesn't necessarily know anything about that, and they don't necessarily care.
If/when the time is reached when external search engines are no longer present in the LLM loop, it seems likely that regular folks won't even notice this shift. As long as the answer-making machine keeps making answers, they won't have any reason to pay attention to this kind of back-end minutiae at all.
It gets its answers from an archive of whatever used to be the Internet, and the last remaining publication: The Costco Flyer.
> How do we double check the answer-machine's answers?
That's the best part: We don't!
and maybe you really need the index and not the engine?
On the LLM side as well, I have not seen much people using Gemini vs the market share of Claude / Open AI.
The assumption of the long tail still using Gemini because it is bundled might be correct, but I am not even sure if that is something that Google will be happy with.
Maybe look into contributing to existing alternatives like PostmarketOS.
Let me revise that super abbreviated prognostication to hint at what I was really trying to say.
I think a lot of people are going to start to look at replacing Android and iOS, and some of these will be decently funded teams. Some perhaps not even based in the US.
Attacking the phone market used to be unthinkable. Not even Facebook could pull it off.
But now? I think it's open season again.
Even if the problems to solve are truly deeper than they appear, LLMs are going to give people so much energy to look past the difficulty.
bear in mind we're on Hacker News. Google's market share is still above 90% - in almost any other market this would be a ridiculous monopoly.
https://nypost.com/2026/07/22/business/reddit-news-outlets-w...
https://www.wsj.com/business/media/google-search-publishers-...
To which, I cannot say with enough emphasis: OUCH. This kills Google Search. All hail Cloudflare Search, the new center of the internet!
No way this is threat to google. For the same reason why the same hordes of engineers did not managed to compete with large companies up to now. And for the same reason they were not producing all that many novel small apps last 10 years.
This is winner takes all economy. Tokens or no tokes, this is not the econony of small competing companies. This is the world of big ones.
Every competitor dies in the womb because "subscriptions are bs and ads are cancer" is totally normalized.
If you want Google to fall, start giving ad-loads or money to companies trying to compete.
But you can't be shit and also charge a subscription. There has to be good stuff behind the subscription.
I don't like Google very much, but making it illegal to scrape public data enables way too much abuse, so this is for the best.
Google vs. SerpApi: The Court Granted Our Motion to Dismiss
But only vaguely. Google uses its monopoly position in advertising to basically ensure that you allow them to scrape your site (or if not you personally, the majority of revenue driving sites). They have the benefit of being allowed by default.
They also then scrape again at the user level for users operating chrome.
They also conveniently ignore global blocks for their adsbots (you have to specifically name them to block them).
If you're not Google, you likely don't have this luxury.
My preference would be that governments force search indexes to be public. The exact mechanisms for this can be debated.