HNHacker News
TopNewBestAskShowJobs

ccgreg

245 karma · joined November 16, 2023

CTO at the Common Crawl Foundation
submissionscomments
ccgreg··on Have I been Flocked? – Check if your license plate is being watched
Well, yeah. Clownflare
ccgreg··on Olmo 3: Charting a path through the model flow to lead open-source AI
Common Crawl is a particular dataset. commoncrawl.org
ccgreg··on Xortran - A PDP-11 Neural Network With Backpropagation in Fortran IV
> it used a stack machine

Do you have some reading for this? I've used that compiler but I never read the resulting assembly language.

ccgreg··on You Don't Need Anubis
Most academic AI research and AI startups find Common Crawl adequate for what they're doing. Common Crawl also has a lot of not-AI usage.
ccgreg··on Synthetic aperture radar autofocus and calibration
I had fun reading this -- the radio astronomy technique called VLBI (very long baseline interferometry) has a ton of overlap, but the jargon words are fairly different. There are plenty of differences: our telescopes are mostly on the surface of the Earth and don't move, VLBI isn't a radar so there's no waveform, etc.

The EHT black hole telescope is an example of VLBI.

ccgreg··on Why I'm not rushing to take sides in the RubyGems fiasco
I've seen a Nobel Laureate disinvited from speaking at a conference due to an incredibly offensive statement. Toxic speech has consequences.
ccgreg··on Athlon 64: How AMD turned the tables on Intel
> they were actually very performant

Insanely expensive for that performance. I was the architect of HPC clusters in that era, and Itanic never made it to the top for price per performance.

Also, having lived through the software stack issues with the first beta chips of Itanic and AMD64 (and MIPS64, but who's counting), AMD64 was way way more stable than the others.

ccgreg··on AI is going great for the blind (2023)
A lot of publishers do not care about blind people, and would prefer that they be unable to use AI to read.
ccgreg··on AI is going great for the blind (2023)
The IETF AI-Preference standard group is currently discussing whether or not to include an example of bypassing AI preferences to support assistive technologies. Oddly enough, many publishers oppose that.
ccgreg··on AI web crawlers are destroying websites in their never-ending content hunger
My experience is that a news crawl is not a big expense at scale, but so far I've only built one and inherited one. BTW No one uses blog pings, the latest hotness is IndexNow.
ccgreg··on AI web crawlers are destroying websites in their never-ending content hunger
That's not how search engines work. They have a good idea of which pages might be frequently updated. That's how "news search" works, and even small startup search engines like blekko had news search.
ccgreg··on AI web crawlers are destroying websites in their never-ending content hunger
I don't think you're correct about Google. Caching webpages is bread-and-butter for search engines, that's how they show snippets.
ccgreg··on AI web crawlers are destroying websites in their never-ending content hunger
The Fastly report[1] has a couple of great quotes that mention Common Crawl's CCBot:

> Our observations also highlight the vital role of open data initiatives like Common Crawl. Unlike commercial crawlers, Common Crawl makes its data freely available to the public, helping create a more inclusive ecosystem for AI research and development. With coverage across 63% of the unique websites crawled by AI bots, substantially higher than most commercial alternatives, it plays a pivotal role in democratizing access to large-scale web data. This open-access model empowers a broader community of researchers and developers to train and improve AI models, fostering more diverse and widespread innovation in the field.

...

> What’s notable is that the top four crawlers (Meta, Google, OpenAI and Claude) seem to prefer Commerce websites. Common Crawl’s CCBot, whose open data set is widely used, has a balanced preference for Commerce, Media & Entertainment and High Tech sectors. Its commercial equivalents Timpibot and Diffbot seem to have a high preference for Media & Entertainment, perhaps to complement what’s available through Common Crawl.

And also there's one final number that isn't in the Fastly report but is in the EL Reg article[2]:

> The Common Crawl Project, which slurps websites to include in a free public dataset designed to prevent duplication of effort and traffic multiplication at the heart of the crawler problem, was a surprisingly-low 0.21 percent.

1: https://learn.fastly.com/rs/025-XKO-469/images/Fastly-Threat...

2: https://www.theregister.com/2025/08/21/ai_crawler_traffic/

ccgreg··on AI web crawlers are destroying websites in their never-ending content hunger
The blekko search engine index was only 1 billion pages, compared to Common Crawl Foundation's crawl of 3 billion webpages per month.
ccgreg··on Stone Age settlement found under the sea in Denmark
Salt water without oxygen and salt water with oxygen are different.
ccgreg··on Cloudflare Radar: AI Insights
It's a similar loophole as public libraries. When I was a kid, I read thousands of books from the library, without paying anyone anything.

But as for the crawl loophole: CCBot obeys robots.txt, and CCBot also preserves all robots.txt and REPL signals so that downstream users can find out if a website intended to block them at crawl time.

ccgreg··on Cloudflare Radar: AI Insights
One way that Cloudflare is gatekeeping is by declaring which bots are AI Bots. Common Crawl's CCBot is used for a lot of stuff -- it's an archive, there are more than 10,000 research papers citing common crawl, mostly not AI -- but Cloudflare deems CCBot to be an "AI Bot", and I suspect most website owners don't have any idea what the list of AI Bots is and how they were chosen.
ccgreg··on Cloudflare Radar: AI Insights
As the Cloudflare post indicates, most crawlers can be verified by IP address.
ccgreg··on The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
Conventional crawlers already have a way to identify themselves, via a json file containing a list of IP addresses. Cloudflare is fully aware of this defacto standard.
ccgreg··on How did .agakhan, .ismaili and .imamat get their own TLDs?
Given that this discussion was started by someone mentioning only these 3 related things, I'd guess that the motivation might a negative one.
ccgreg··on VLT observations of interstellar comet 3I/ATLAS II
You can leave off the laser beams and just look at them with telescopes. The paper we're discussing is exactly that.
ccgreg··on VLT observations of interstellar comet 3I/ATLAS II
Undetected means you can compute an upper bound, based on the performance of your instrument. Again, these are undergraduate level concepts.
ccgreg··on VLT observations of interstellar comet 3I/ATLAS II
No. Undetected iron and zero iron are different. Source: astronomer.
ccgreg··on VLT observations of interstellar comet 3I/ATLAS II
> We report detection of CN emission and also detect numerous Ni I lines while Fe I remains undetected, potentially implying efficiently released gas-phase Ni.

Where does it say there is zero iron? This is an upper bound, not zero.

ccgreg··on Lightning declines over shipping lanes following regulation of sulfur emissions
https://en.wikipedia.org/wiki/Global_dimming

Sadly misunderstood by a bunch of people.

ccgreg··on Lightning declines over shipping lanes following regulation of sulfur emissions
Hopefully you read all of the links in the article -- the purpose of thecoversation is to present information to the general public, with references to research that the author has been involved with.
ccgreg··on Show HN: Building a web search engine from scratch with 3B neural embeddings
> Publishing a crawl, or the URL's, under CC-0, CC-by, BSD, or Apache would make them usable without restrictions or any further legal analyses.

This isn't true, and I can't imagine that any lawyer would agree with this statement. CCF does not have rights ownership of any of the bytes of our crawl, so we cannot grant you any rights for the bytes in our crawl. Nothing that we could say could have any relationship to this legal issue.

ccgreg··on Non-Uniform Memory Access (NUMA) is reshaping microservice placement
Historically, the Sparc 6400 was derided for not being NUMA, but instead being Uniformly Slow.
ccgreg··on Non-Uniform Memory Access (NUMA) is reshaping microservice placement
> The long and short of it is that if you’re building a HPC application, or are sensitive to throughput and latency on your cutting-edge/high-traffic system design, then you need to manually pin your workloads for optimal performance.

Last time I was architect of a network chip, 21 years ago, our library did that for the user. For workloads that use threads that consume entire cores, it's a solved problem.

I'd guess that the workload you had in mind doesn't have that property.

ccgreg··on ArchiveTeam has finished archiving all goo.gl short links
The goo.gl URLs that are publicly known are already in the Internet Archive and Common Crawl crawls.
← PreviousPage 3 of 5Next →