I ask because a lot of what you discuss seems to resonate with folks in the graphics business who are, for example, culling object geometry from a render scene by doing what is effectively the above operation. Its also what a 'whisper' or 'yik yak' type of App does trying to find people near itself although in that case the problem is greatly simplified by assuming a sphere and testing for containment in the sphere (or the circle if being planar)
The geospatial implementation in SpaceCurve is quite sophisticated and in some aspects the state-of-the-art. Complex geometry types and operators are obviously supported, and the implementation is fully geodetic and very high precision by default. Complex polygon constraints, aggregates, joins, intersections, etc on data models that also contain complex geometries are supported and massively parallelized. There is nothing else comparable for big data in this regard that I know of.
It is useful to understand that the entire platform, database engine on up, is fundamentally spatially organized even for plain old SQL "text and numbers" data. The SQL implementation is a thin layer that is translated into the underlying spatial algebra. You can use it for non-geospatial things, it just lends itself uniquely well to geospatial data models because it can properly represent complex spatial relationships at scale.
Have I got the wrong impression?
There are an unbounded number of possible shapes and surfaces in 3-space. Storing and content-addressing them is easy enough. The challenge is that to be useful you need, at a minimum, generally correct intersection algorithms for all of the representable shapes. It is an open-ended computational geometry problem and we continuously extend geometry capabilities as needed. Currently, the most complex shapes can be constructed relative to a number of well-behaved mathematical surfaces embedded in 3-space. Shapes directly constructible in 3-space are significantly more limited but I expect that to expand over time as well, particularly if there are specific requirements and use cases.
The underlying representation was intended to mirror the spatiotemporal organization of the physical world at a data level. We can generate projections but the data lives in 3-space.
This included information such as land elevation, floods areas, soil types, distances to roads and their importance (highway, major motorway, how many lanes), distance to facilities (electricity / gas / water / sewer), population demographics (wealth, buses) etc.
So beyond just measuring distances and polygon coverage, your GIS will have a lot of layers of data. Each layer either representing different data in itself, or just more detail. The query then is to find best matches knowing data, criteria, and rankings.
On a related note, I was tremendously impressed with the quality of discussion at imechanica.org. At a glance, it may be the best example of a well-functioning online community that I've ever seen.
I'm having trouble thinking up examples - can you provide some?
Some of the sexier petabyte-per-day sources include mobile phone telemetry, vehicle telemetry, and countless remote-sensing networks and platforms (of which there are far more than most people imagine). A petabyte per day is not that much data per entity, there are a lot of entities, and you can support this data rate in a single rack of commodity hardware.
A more typical generic IoT application at a large company is 1-100 TB/day, but that is quickly creeping upward as it becomes more cost effective to deploy sensors and store/analyze the data. And more often than not, they need to store weeks or months of that data so the overall data set size is still quite large.
Can you share company names or more concrete examples?
Even ups only delivers 17M packages / day.
edit: also, which industries customarily have thousand vertex polygons? Thanks for the interesting blog post.
The interesting time series are heavy machinery telemetry feeds (for example a Rolls Royce jet engine has its own communications subsystem to phone home to a RR engineering station via satellite, which however, likely makes it low-bandwidth).
Knowing where everyone is at all times; the ultimate enabler for weaponised harrasment.
Or if your the NSA, tracking mobile phones.
Searching for a generic/theoretically all-encompassing solution is quite akin to looking for the perfect pub-sub system for all workloads. :)
My comment was related to lat/lon indexed data clustered around cities. If you know beforehand that you're going to deal with such a dataset (static or dynamic), one can pragmatically decide on some sort of by-convention partitioning (e.g. Partitioning by continent/region which will bring it down to in-memory indices not needing continuous disk access).
I love the non-euclidean bit in the piece you wrote. Anyone who has tried to do a k-nearest neighbor query east of New Zealand will appreciate that bit. :D
- The average distribution of the data and the instantaneous distribution of the data can be very, very different. This means that some cells are overloaded while others are idle, and the whole system runs as slow as the overloaded cell. The canonical (and fairly benign) example is data following the sun.
- Many data sources have inherently unpredictable data distributions.
- Spatial joins across different data sources (say, weather and social media) require congruent partitioning or it won't scale. Unrelated data sources tend have unrelated data distributions, so this is a problem.
Making your partitions match your data distribution is good practice for static data layers. With some caveats, this can be modified for spatial joins across static data layers as well. For dynamic data sources, you run into issues with data and load skew at scale.
everyone who built any online service know the access pattern is a wave with time of day.
do you mean to say you've seen people trying to shard a geodb by latitude instead of longitude? ... that would be a very sloppy initial research.
or i guess the sun is not predictable? sigh...
See my comment: https://news.ycombinator.com/item?id=9208066