S2: Computational geometry and spatial indexing on the sphere
github.com
github.com
I find it liberating to not rely on a backend implementation of geospatial indexing, or rather being agnostic to backend, simply indexing very efficiently the 64-bit integers.
The hierarchical nature of the Hilbert curves means that the child cells contain the ID of the parent, and also fit perfectly into the parent. This is in contrast with H3 where you necessarily cannot fit the 7 child hexagons into the parent without some overlap (see for example the logo on the H3 website https://eng.uber.com/h3/)
What is an official Google product? What does it mean to be one? What should I understand from the presence or absence of this disclaimer? Probably some sort of SLO, but what precisely?
Tensorflow doesn't have this disclaimer: https://github.com/tensorflow/tensorflow
That it will eventually get killed.
> What should I understand from the presence or absence of this disclaimer?
That it will probably last longer. Open-sourcing helps here too obviously.
Anyone knows why they used a sphere model and why they didn't just use GeographicLib (MIT license)?
ruby:
S2Cells::S2LatLon.new(34.1234, -84.1234).to_s2_id(30)
=> -8577780315821624425
S2Cells::S2LatLon.new(34.1234, -84.1234).to_s2_id(1)
=> -8358680908399640576
https://gojekfarm.github.io/s2-calc/ agrees with these values
python's s2sphere (from sidewalk labs, seems like it'd be more of a reference implementation) gives very different numbers:
s2sphere.Cell.from_lat_lng(s2sphere.LatLng.from_degrees(34.1234, -84.1234))
=> Cell: face 4, level 30, orientation 0, id 9868963757887927191
This discrepancy feels like it'd be some kind of signed/unsigned integer issue, but I didn't think that would be a thing in Ruby. It also doesn't seem like the level=1 cell value in Ruby would be such a large, specific number. I'm missing something.
Level 1 IDs look like large random numbers in decimal, but when viewed in binary they have many trailing zeros, that are used to truncate to make the tokens.
-8577780315821624425
and
9868963757887927191
have the exact same underlying 64-bit representation, which is all that matters for S2 (you can verify this by checking their difference is exactly 2^64). The former is just interpreting the 64-bit string as a signed int, while the latter is unsigned.
Separately, the level 1 number you mention
-8358680908399640576
is fairly simple if you look at its binary representation; almost all of its trailing digits are 0.
Basically, you really shouldn’t think of S2 values as integers, but more as their underlying length 64 bit string.
Secondly, GDAL is tightly bound to Proj [1] for reprojecting spatial coordinates data from one spatial reference system (SRS) to another. The "sphere" libraries imply no generalized spatial reference, instead only a single spherical one.
Basically, you shouldn't ever be choosing between the two. If you're thinking of using a representation system like S2 for raster data (e.g. images or anything else on a regular grid), rethink things a bit.
But your point about raster data is correct. If you're given a raster, like a satellite image, S2 has no built-in type to deal with it.
On the other hand, if you can define your own raster, like if you're making a heatmap or something and don't care about it being a lat-long aligned grid, you can use S2 cells like a raster, as they're roughly square and they tile the surface.
This is pretty common as it's fast and convenient. It has some advantages over lat-lng grids as well, because S2 cells are roughly equal in size across the whole surface of the sphere, while lat-long aligned pixels obviously get really warped near the poles, etc.
You're basically making a row in a DB for every _pixel_ with the S2 approach.
The main point of raster storage is that it avoids the overhead associated with the "row for every datapoint" approach. E.g. the X and Y are implicit and it's simply a big array of Z values, then (optionally but commonly) a sequence of additional downsampled Z arrays for efficient lookup of low-res versions. These are typically tiled for efficiency of extraction of sub-regions. The simplest systems are flat pyramids of raster files in a bucket or on a filesystem. Things like tileDB are essentially a couple of layers on top of this type of idea. Either way, the basic unit is a few thousand pixels instead of 1 pixel, which is generally more efficient, as users are usually requesting millions of pixels.
A key advantage is that things are fundamentally in raster format and don't need to be translated back to raster format for the end user.
Basically, the requests / etc that wind up being made are for fundamentally pixels that are parallelograms of some sort in regions that are parallelograms of some sort. After all, this has to go into a .tif/.png/.hdf/etc container at the end of the day. There are a lot of advantages to storing data in a form that's close to that.
(I work for Google, have worked in Google Geo for 10+ years and I'm on the Geometry team now, but I'm not speaking for them officially of course)
I don't know enough about GDAL to say anything about the libraries relative abilities.
Last big discussion 2018: https://news.ycombinator.com/item?id=17849546