There are a couple issues I see with basic access and working with large datasets. Ease of access for typical users is also a valid issue.
First, we still mostly move the data to the computation when we should be moving the computation to the data. Moving the data works fine when data is small but if the data volumes are large (as sensor/geo data tends to be) then it can take an incredibly long time to move the data. In many cases, more time is spent shoveling data over the network than actually doing the computation. This has become worse as storage density has increased, hundreds of TB/server is ordinary.
Second, the data is rarely organized in a way that makes it efficient to extract arbitrary subsets. There is still a lot of what is essentially "grep at scale" going on. Again, not a problem if the data is small but if I need a specific 50TB subset of a 10PB source, this becomes prohibitively slow. The data needs to be organized such that we can slice and dice it with high selectivity in place, much more like a proper database and less of a distributed filesystem. Because spatiotemporal analysis tends to involve iterative join-like operations, you want this to be efficient as possible.
The other big problem is many of these data sources are too large for everyone to have their own copy. Or if they did have their own copy, it would be extraordinarily wasteful. This is adjacent to the first issue. EDIT: And herein is the likely business model.