You can usually check the ads.txt file on a website to see which companies are allowed to bid for ad space on there. For example, for dict.cc, the website in question:
The ones labelled "RESELLER" will probably share your data with even more ad companies.
It generally goes like this:
When we launch a site it is seldom more than perhaps Hotjar, Google Analytics, and two-three other services connected.
And then through the years product managers and other stakeholders gets sold on adding LinkedIn, Instagram, Meta, and so on. So we add those.
Next a specific service ”to better track the sales funnel from in-store salespeople to the web” gets added. Then another ”analyse the data quality versus bounce rate” tracker gets added. And so on.
Before long the developers have streamlined the process of adding new scripts/analytics/trackers that editors can add them on their own, and that is when the floodgates open.
analytics: A/B testing, "if x does user click y"?, unique page visits, etc.
ads: integrating with an ad provider comes with hundreds of trackers, because they want to - know if you bought a product after clicking on an ad - show you targeted ads for shoes after you googled shoes - build a profile of you (age, gender, location, profession) to show relevant ads across different websites
For most companies this can easily be thousands of partners, and going through that list and figuring out exactly who might get data in reality, through every possible permutation of workflow, is a horrendously expensive proposition.
You might be surprised how many well-meaning regulations leave even the best-intentioned implementers in an impossible situation.
And once again we shall see how being conservative sounds like it might save you money but costs you dearly in the long run.
Ultimately that is what they are having to do though, it's just costing them twice as much by pretending that being conservative and not actually looking at the problem saved them.