I think it depends on whether you are talking ad-hoc ie. a user analyzing several datasets, or pipelined ie. preprocessed joins.
For pipelined joins, effectively your data forms a DAG (directed acyclic graph, and yes we are ignoring recursion here). Providing your data services speak the same language you can create a pipeline off the first pipeline that joins the data and sticks it in a cache (eg. RDS, Elasticsearch).
Changing the underlying data should then trigger a reload of the downstream pipelines.
This is basically what Materialize.io, KSQLDB et. al. do - a reactive DAG with a database as the cache.
One issue for larger companies is that you don't control the whole DAG, so discovery, security, protocols etc. need to be coordinated by an overarching architecture for this to work.
Something like Apollo (GraphQL) is a simpler solution (in some ways) as you control the joins on the Apollo server which speak to backend (REST) APIs (other teams).