Geospatial data science with Julia
juliaearth.github.io
juliaearth.github.io
Given these circumstances, how might the incorporation of Julia and some geospatial DB (PostGIS) contribute to further optimizing geospatial data retrieval and presentation, especially when dealing with large datasets and intricate geospatial operations?
I don't know Julia well, but I definitely would suggest exploring whether PostGIS can help improve the speed of your DB queries.
I'd also consider how you deliver your geospatial data to your clients -- I'm not sure GeoJSON is your best bet. Protobuf tiles might be better for your use-case (e.g. the Mapbox Vector Tiles spec).
For anyone working with GIS data, it's absolutely worth investigating what PostGIS provides and the ease of integration to your existing application!
PostGIS gives you the benefit of spatial indexes which are extremely performant.
I've seen Python GeoSpatial applications taking hours to finish processing which only took a few minutes when shifted onto PostGIS.
If you're also doing a lot of processing in Python, exploring other languages could also help. In the case of Julia you get a typed language that's also JIT compiled.
https://geopandas.org/en/stable/docs/reference/sindex.html
I think that the challenge for most is that the PostGIS query planner does the indexing for you in most queries, while a naive all-pairs comparison in geopandas/shapely won't tell you to use the .sindex attribute instead.
https://osmfoundation.org/wiki/Licence/Attribution_Guideline...
[1]: https://juliaearth.github.io/GeoStatsDocs/stable/domains.htm...
[2]: https://geopandas.org/en/stable/docs/user_guide/set_operatio...
https://github.com/JuliaGeometry/Meshes.jl https://github.com/JuliaGeometry/Rotations.jl https://github.com/JuliaEarth/GeoStatsBase.jl https://github.com/JuliaEarth/PointPatterns.jl
- Generate high-performance code
- Specialize on multiple arguments
- Evaluate code interactively
- Exploit parallel hardware
> This list of requirements eliminates Python, R and other mainstream languages used for data science.
Can you elaborate on why/how? Awesome work by the way
It should be noted that this is usually sufficient. But particularly for earth scale problems it can often not be.
If so, it looks like you're interfacing from R to high-performing code written in C. Isn't that exactly what OP was describing?
Now, I generally use Julia for heavy computes, and usually its much faster than R. But not always.
And this little bit of code runs for hours on the largest instance on AWS every day. Why I was looking so speed it up.
Like, R is “what if we made a lisp inspired version of Python built around numpy and pandas and then reversed timed”
And data.tables in R is faster (and I think nicer to write) than DataFrames in Julia. And since data.tables feed my optimization, R still wins.