HNHacker News
TopNewBestAskShowJobs

Mosarwat

56 karma · joined March 18, 2022

submissionscomments
Mosarwat··on Apache Iceberg now supports geospatial data types natively
Please note that not all query engines supports the native geo type in Iceberg yet. The first one to support it is Apache Sedona, which works well with Spark:

https://github.com/apache/sedona

However, the ultimate goal is to make more engines (e.g., Arrow, Trino...) support the geo type too

Mosarwat··on Apache Iceberg now supports geospatial data types natively
Beside Zarr, there are also efforts to support different types of raster (sort of multidimensional) data such as geotiff and NetCDF. The Iceberg Geo spec was heavily influenced by the Havasu project proposed by Wherobots, which also supports that type of raster data. However, the Iceberg geo spec still only supports only geometry for now.

https://wherobots.com/building-a-spatial-data-lakehouse/

Mosarwat··on Apache Iceberg now supports geospatial data types natively
Now, Apache Sedona is the first engine that will support that native geo type in parquet, but Arrow will also support it very soon.
Mosarwat··on Apache Iceberg now supports geospatial data types natively
GeoParquet will still be used for bit since both the native geo types in parquet and geoparquet have slight differences. But, the ultimate goal is to shift all data / workloads to the parquet native geo in a couple of years!
Mosarwat··on Apache Sedona: Big Geospatial Data and AI Engine
Apache Sedona™ is a cluster computing system for processing large-scale spatial data. Sedona equips cluster computing systems such as Apache Spark and Apache Flink with a set of out-of-the-box distributed Spatial Datasets and Spatial SQL that efficiently load, process, and analyze large-scale spatial data across machines.
Mosarwat··on The Apache Software Foundation Announces New Top-Level Project Apache Sedona
The geospatial analytics market is expected to reach $300B by 2030 (Source: Precedence Research). Apache Sedona provides the technology that enables companies to process data at scale, realizing tremendous business value. Users and contributors of Apache Sedona include some of the world's largest e-commerce, telecommunications and data management companies.

Apache Sedona provides APIs and libraries for working with geospatial data in the Scala, Java, Python and SQL programming languages, and it offers support for a wide range of geospatial data formats. It also provides tools for spatial indexing, querying, and spatial join operations, as well as support for common spatial analytics tasks such as clustering and classification. Apache Sedona is designed to enable scalable and efficient analysis of large datasets, and it can be deployed in standalone, local, or cluster modes.

“It has been incredibly exciting to see Sedona empower so many users to gain new insights from their geospatial data,” said Jia Yu, Apache Sedona PMC Chair. “We are also excited to add new features and capabilities for Sedona including support for more data processing engines and integration with popular geospatial map visualization and GIS tools.”

Key Features of Apache Sedona include:

Support for a wide range of geospatial data formats, including GeoJSON, WKT, and ESRI Shapefile; Scalable distributed processing of large datasets; Tools for spatial indexing, spatial querying, and spatial join operations; Support for common spatial analytics tasks, such as clustering, classification, and regression analysis; Integration with popular big data tools, such as Apache Spark, Apache Hadoop, Apache Hive, and Apache Flink for data storage and querying; A user-friendly API for working with geospatial data in the Scala and Java programming languages; and Flexible deployment options, including standalone, local, and cluster modes.

ADDITIONAL RESOURCES

Website: https://sedona.apache.org/ Github: https://github.com/apache/sedona Twitter: https://twitter.com/ApacheSedona Discord: https://discord.gg/GUbGJGdK

Apache Sedona is steadily growing with nearly 100 committers to the project and 800,000 monthly downloads. It also ranks among the top 1% most downloaded Python projects on PyPi.

Mosarwat··on Apache Sedona for Processing Geospatial Data at Scale
More details at the system here: https://sedona.apache.org

New about it on Twitter: https://twitter.com/apachesedona

Mosarwat··on Apache Sedona for Processing Geospatial Data at Scale
Apache Sedona is an OSS for processing geospatial data at scale. Sedona provides a turnkey way to:

1. Efficiently ETL and Integrate geospatial data from heterogeneous sources

2. Query, Process, and Analyze Geospatial data at any scale

Sedona's benefit is two-fold: (1) Easy to Use API (using standard/popular spatial language) and (2) Fast & Efficient processing (100X faster than geospatial tools and 10X faster (with 50% better resource utilization) than other big data systems Implementations)