If I want a worldwide ad (is that even a thing?) for left-leaning horse owners between ages of 20-27.5 with a child and at least 2 partners, can’t that be dished out from an EU server for EU users?
If I want a worldwide ad (is that even a thing?) for left-leaning horse owners between ages of 20-27.5 with a child and at least 2 partners, can’t that be dished out from an EU server for EU users?
It is not clear what data they are worried about transmission of, but each type seems to need special consideration. You mention ad-targeting data, but most data collected by Facebook useful to the surveillance agencies Ireland is worried about are more personal than that.
The question is potentially very difficult, and could only be resolved by constructive engagement with the party making the rules.
Don't forget what happens when people travel.
Borders are heavily controlled by USA, why shouldn't other countries to the same?
We do want to make it easier for Facebook for political reasons, though, and it's still not particularly hard: just declare that only the data/conversations involving US citizens can be stored on US servers.
Knowing people in other countries is a tiny minority? Maybe in the US, given its size, but it's pretty widespread and normal in Europe.
Absolutely yes.
Especially across two continents.
Europeans who have friends from Europe would all be in Europe anyway.
Europe is larger than US btw, it has two times the population.
Russia+Turkey+Germany alone account for 95% of the population of the United States.
Globally, around 1 person in 30 lives outside their country of origin, so knowing people in other countries should be common.
Not most, just some.
1 out of 30 is a bit more than 3% and many of those have family connections, they are not strangers living abroad, they are - for example - Italians living in Canada.
EU citizens whose data should be kept in EU.
I believe, in the context of this proposed ban "Europe" means "the EU", not geographical Europe. So both Russia and Turkey don't count.
I belong to a group that has several connections in countries all around Europe due to frequent traveling to dance events. Most of these people post on Facebook in their own language and attend local events.
Outside of this bubble, things are very different. The majority of people never move from where they are born, speak poor English and never travel.
While the people with international connections are surely a relevant amount, and even adding expats that keep contact with friends and family, it's the group with not international connections that I would define as "widespread" and "normal".
With Facebook’s lowering importance and my self getting older, my interaction with non Americans has lowered. If we are talking about the EU only, the amount is minimal.
Thinking about others around me, most don’t have anything significant with EU residents or have one specific set of friendship[s] in the EU.
Sounds almost philosophical. I have a friend. We both have brains. In which brain is our friendship stored?
If we're feeling especially poetic, we could make the case that friendships can live on despite the death of one of the participants.
Also, do we always assume a binary friendship of exactly two participants?
https://www.goodreads.com/quotes/548471-people-who-live-in-s...
Let us consider your hypothetical glitched friend-state within Antoine’s concept of distributed self awareness. Does the friend delusion lead to self delusion or vice versa?
EU GDPR is pretty clear that all of the cases which involve transferring data from the EU to the US fall under EU GDPR rules. It refers to the location of the data, not the citizenship or residency status of the individuals who are involved.
So, the photo that you took and uploaded within EU borders is theoretically in scope for EU GDPR - even if your French friend is not in the photo.
Otherwise, a lot of this data presumably isn’t covered by GDPR. You and your friend are ultimately user_ids with a relationship_id or sth, and these surely aren’t GDPR-controlled (but they point at details that are)
Your message will be stored in the US. Your friends message in the EU.
The metadata ( conversation table of both) would be stored in the US.
Your picture from Europe as a US user with US residency will be stored in the US. As Facebook wouldn't want your data under GDPR and since it's also not required.
If you have a full datacenter (i.e. containing databases, not just a frontend or CDN / PoP footprint) in a country, then typically the entire logical data set -- all data for all users worldwide -- is presumably replicated there. Other systems and services will then make assumptions that any object can be looked up with low sub-ms latency.
Social networks often contain activity streams and other pages that include content from many users at once. Consider algorithmic ranking of feed content, comments on popular page content, etc: how do you even implement this if even some small subset of the data needs to be fetched from halfway around the world on every page view?
Anyway, to answer your original question, friendships are bidirectional associations and would typically be stored in two places: one entry in your db shard, and one entry in your friend's shard. Photos are objects and presumably would be "owned" by a single user or page and located there (at least in terms of the metadata about the photo); however tags may be associations which have entries on multiple shards just like friendships. If some of these shards can only be accessed across a trans-Atlantic link, the entire scheme falls apart due to the latency.
"The social graph is tightly interconnected; it is not possible to group users so that cross-partition requests are rare. This means that each TAO follower must be local to a tier of databases holding a complete multi-petabyte copy of the social graph. It would be prohibitively expensive to provide full replicas in every data center. Our solution to this problem is to choose data center locations that are clustered into only a few regions, where the intra-region latency is small (typically less than 1 millisecond). It is then sufficient to store one complete copy of the social graph per region. Figure 2 shows the overall architecture of the master/slave TAO system."
https://www.usenix.org/system/files/conference/atc13/atc13-b...
edit: Oh hah, didn't realize parent was former DB person at FB. I was too, just a few years before :-)
Every proper Microservice zoo also has the exact same problem of ownership of data and latency. Holding duplicates of often needed parts of the data works surprisingly well and would also be fine with GDPR.
Option A) client sends feed request to all regulatory domains and mixes results itself. Downside is either you'd always see the tail content from all regions, or client would have fetched content it doesn't display. Also, if client has to contact a load balancer in the regulatory domain directly, you're at the mercy of the client's international transit which is often worse than FB's.
Option B), like option A, but client sends feed request to nearest server, nearest server bifurcates the request to one per regulatory domain. If regulatory domain is local, satisfy request (with all the sharded queries), otherwise send it to a request processor with local data and do the processing there. Feed request processor mixes the data. Downside, extra memory used holding results waiting for the remote regions to respond and managing more requests in progress. You could get tricky and use recent response times to try to get all responses back around the same time, and reduce the time where you had some region's data in memory, but not all of them, but that's probably silly.
Both of these are more work than the current scheme, and nobody likes regulation requiring more work, but still. Figuring out where things are permitted to be stored when they involve people in multiple domains sounds like a headache though. And actually splitting the data once the parameters are clear doesn't sound like my kind of fun either.
Disclosure: I worked at FB, but not with FB user data.
Viewing a couple pages of my news feed randomly just now, I see content from a mix of friends in the U.S., Ireland, Germany, the U.K., and Hong Kong. The content is algorithmically ranked and sorted, and then paginated. Similar situation with viewing comments on popular pages, which may span many countries.
If ranking/sorting is performed either client-side (your option A) or even in a nearby aggregation server (B), each and every user-facing pageload of N posts/comments/whatever would need to fetch the worst-case of N entries from each regulatory domain, and then re-rank/sort/paginate the aggregation of them.
This presents numerous significant problems: not just network latency on every single page load (which would be severe), but also bandwidth consumption, memory, and cpu wherever the aggregation/ranking are being performed. I don't think your option A is even feasible for a worldwide social network, with many users in countries where less powerful devices are prevalent. Even option B would have scary implications for the worldwide increase in network bandwidth consumption alone.
And that's not even considering the problem of how you migrate one of the largest data sets in human history from a non-geo-sharded topology to a geo-sharded one.
Worst-case, multiply all of this effort (and bandwidth impact) by every single multi-national social network, publishing platform, and user-generated content site/app/product. The GDPR-compliance efforts of 2017-2018, which were quite substantial at all of these companies, would look like peanuts in comparison.
Also, let's not forget that we're talking about a service with absolutely horrible reliability. Facebook only cares about reliability at scale; at a single person level reliability is quite poor; you can suddenly discover that your random post has 60 thousand likes (for a few minutes), or a friend of yours has a new post (they don't). Thus, the questions of "what if something has >10ms latency" don't really matter - Facebook fails much worse than that all the time.
Do you believe Facebook is the only company that would ever be affected by this?
Personally I strongly suspect that once the precedent is set -- that is, regulation requiring user data to remain strictly on servers in the user's country and nowhere else -- the entire tech industry would be severely impacted.
> the questions of "what if something has >10ms latency" don't really matter
Internal point lookups for each individual piece of content would go from sub-ms (intra-region) to now being upwards of 200ms (network requests spanning the globe). And for many social applications, a given page contains content from dozens of sources. Much of this could be parallelized, but even then, the total wall time for each parallelized multi-get is determined by whichever source had the worst-case latency. In any case, we're not just talking about tens of ms of impact here, but rather a latency increase of 2 orders of magnitude (or more!) for even the simplest lookup operations.
Of course not, Facebook is not the only company parasitizing on users' privacy. The point still stands, though: it's an absolutely tiny (compared to the company's income) amount of work, versus users's rights.
As for delays - again, FB routinely displays much worse problems, adding some more delay there wouldn't be noticeable.
It's quite clear that GDPR allows data to cross the border, if particular safeguards are guaranteed by the data recipient.
That continues to be the advice generally given following the rejection of "Privacy Shield" e.g. by https://gdpr-info.eu/issues/third-countries/
This particular ruling affects only Facebook (and is specifically against Facebook), because it is only Facebook that has here been identified as failing to offer those safeguards.
Google similarly is a tech giant operating across this border, and doesn't currently face such a sanction.
the us db can have a list of friends by user id and those friends might be in other countries db. however your identity and data live in us db only.
if you create a message the message can either be stored in the us and referenced by id in france (or duplicated since you intentionally sent it there). same for photos
this makes each transmission of data across regional boundaries more intentional and easy to add governance checks
Search "privacy shield invalidated" for more info.
Facebook has the deep pockets to either build this through or fight it legally to the bitter end.
But for any startup trying to build a global X, this can put a serious blocker in the way. Global sharding of entities and working with that isn‘t for the faint of heart. I know we‘d stand still for a year or two trying to implement something similar as a small to mid sized startup.
Philosophically and as a user, I actually like the idea, but wearing the systems architect here, this requirement scares me to no end and could throw a literal wrench into the operations of any global effort.
Social connections span legal jurisdictions.
Easier than CSS.
It's really hard to design systems split like that. There are so many corner cases. What when a user from the EU shares an image with a US group containing users from Australia. Where will the image be sent? Where will it be stored? What happens if the US members leave that group? Need the image be rehomed? What if at the moment that happens there is a trans oceanic bandwidth shortage/outage? Will there be a queue for rehoming images triggered by the final member of a group from a continent leaving the group?
To avoid all the complexity, 99% of 'distributed' systems have a 'master' for the data in just one place. I'm not sure there are any distributed datastores that are masterless outside academia.
More practically through, all of facebook's infrastructure is built on a transparently globally replicated database. Data siloing was only originally considered for launching facebook in China
With regards to targeting: yes, you can target any country (and all of them).
With regards to reach: yes, because you'll reach the user on pretty much any site/app with ads, so it's virtually "running worldwide" (from the user's point of view).
It's devastating to their "cost of doing business" not "conduct of business".