Kart: DVC for geospatial and tabular data. Git for GIS
kartproject.org
kartproject.org
* Our CTO Rob Coup presenting on Kart at FOSS4G 23: https://www.youtube.com/watch?v=1B-HB2Z3Vlc
* Docs are available at https://docs.kartproject.org/en/latest/
* We also have a QGIS plugin! This gives you visual diffs of vector feature changes. https://plugins.qgis.org/plugins/kart/
Happy to answer any questions!OSM got vandalized recently [0][1] and apparently the community is having a hard time restoring the data / reverting the edits.
Would using KartProject enable storing edits as commits which could then get easily reverted? Also localized commits?
[0] https://www.openstreetmap.org/#map=17/32.09438/34.77448
[1] https://www.openstreetmap.org/#map=16/32.0914/34.7746&layers...
That said, the general issue of "data supply chains" and keeping track of who did what, and when, is largely unsovled outside of OSM, and if you're looking after your own GIS data and you're concerned with tracking and detecting unwanted changes, Kart is a great option. You get all the cryptographic verifyability of git for free.
1. We found Geogig hard to get running. I'm a software engineer but I struggled to get Geogig installed correctly. Typical GIS people are technical, but maybe not in command lines or building software. Kart is designed to be 'batteries included', installs easily on Windows, Mac and Linux.
2. Geogig was "inspired by" git, but isn't built "on top of" git. That meant rebuilding a lot of things that git is already _really_ good at. Kart teaches git how to work with spatial data, so we're getting all of the benefits of many thousands of human-hours in optimizations in git. It also means a lot of git tooling works with kart today, leaving us free to focus on GIS-specific enhancements. There are a few specific examples of later additions to git, like git filters, that have made critical features possible (e.g. spatially filtered clones).
3. Geogig was a great project largely sponsored by Boundless Geo before they got bought by Planet, and it more-or-less died at that point. Ultimately, these projects have a big bootstrapping problem. You need data, the tool, and willing users. Kart is sponsored by Koordinates.com - a platform that's been doing GIS data delivery for over 10 years. We have a lot of data that we're been 'mirroring' into Repos and will make avaiable soon, we have a lot of users with specific use cases (fetching updates to large, regularly updated datasets) who are already using it, and we're making a long term committement to Kart as an OS project.
My project never took off because it was my first cloud and Maven project, but it was fun to tinker with the idea until GitHub got around to its GeoJSON stuff.
https://docs.kartproject.org/en/latest/pages/development/tab...
Top takeaway being that it's not just versioned geo feature items, but versioned per-feature formats. Various popular GIS database formats are supported as the "checked-out" representation, analogous to a git local-filesystem tree. Maybe does conversions between standard GIS formats well -- wasn't obvious.
One question I'm left with is performance:
> Every database table row is stored in its own file. ...
So, we're seening pretty good performance. We're maintaining a number of repositories with several millions features, with a decade of weekly updates of ~10,000+ rows. It _does_ take some time to push that data around, but it's _vastly_ better than old ways, and once you have your clone, maintaining updates becomes extremely trivial - a _major_ unsolved problem in the GIS/data world.
I'd add - Kart has GIS specific features that nullify some of these issues. The ability to spatially index the objects, then filtering them on Clone, means I rapidly clone a tiny subset of the data to work with.
Okay -- so is the "--depth=N" filtering option to git-clone supported as well? And does it remain useful in the context of Kart applications?
Are the raw files in the working repository GeoPackages? How is it tracking the changes made inside the geopackages? What happens if it's replaced with an updated copy of the geopackage the was edited via some other application? How does it diff the changes?
> Are the raw files in the working repository GeoPackages?
The working copy for a vector/table dataset can be in a GeoPackage or a SQL database like PostGIS. For rasters/point-clouds they're flat files.
> How is it tracking the changes made inside the geopackages?
In general, triggers which store RowIDs/PKs of inserts/updates/deletes. Then when you ask for a diff or make a commit Kart figures out any actual row-level (or schema) differences.
> What happens if it's replaced with an updated copy of the geopackage the was edited via some other application?
If it's edited by something else (QGIS, ArcGIS, python/go/whatever application, SQL CLI, whatever) it'll work: you do edits where you want to. If it's replaced by something else, it won't work.
> How does it diff the changes?
Comparing the features/rows in the repository (and their schemas) against the rows in the working copy database. It uses the stored list of modified rowids to make this fast.
I don't see how they'd display anything other than points. That leaves XYZM for diffing.
They might be showing summary stats like length, perimeter, area, volume, but that's usually not easy to generalize.
Kart supports points, lines & polygons, as well as GeoTIFFs for imagery and LAZ for point clouds.
Kart is a CLI tool, but provides fully machine readable outputs. You can use the QGIS plugin to get a visual diff of vector feature changes though.
This sounds like an amazing opportunity for a screenshot, btw :)
Does this improve Git's support for large binaries generally, or is it necessary to have introspection into any filetype you want to support?
Is there good interoperability with existing Git repos?
> Because Kart uses Git for data transfer and storage, you can host a Kart repository anywhere you can host a Git repository - for example, GitHub, Bitbucket...
ref: https://docs.kartproject.org/en/latest/pages/basic_usage_tut...
No, this still uses LFS for larger binary formats (ie raster or point cloud datasets)
Kart serializes vector/tablular data into datasets in the repository, and manages the process of writing them out to useful working copies (GeoPackages, or into databases).
For large binaries - rasters and pointclouds - we're using LFS. We include some additional spaital information into pointer files to enable some very useful GIS functionality, like spatially filtered clones (this works for vector data too).
OpenStreetMap has it's own versioning mechanisms (and a fairly specific-to-OSM data model) and Kart isn't really designed to work with OSM data as such. Kart adds version control to the GIS data that planners, academics, architects, civil engineers, etc, use day-to-day. There's a lot of data out there!
"Large" is relative, but Kart works well with quite big vector datasets for these typical use cases. For example, we're regularly working with datasets that have over 2 million features, with a decade of weekly data changes.
Kart includes some feautres specifically for working with small geographic areas. We can spatially filtering cloned data so you're working with a small subset of a much larger dataset, but you still retain the abilityt commit/merge/push to the source repo.
Kart gives you row-level tracking, so you can see who made what change & when, and diffs small and fast to apply.
Kart works with GIS working copies that are more familiar to GIS people - e.g. GeoPackage, Postgres/PostGIS & MSSQL databases. Differenet users can use different working copies, and still collaborate together too.
* Docs should not be hidden in small font and as disabled link color, make it big button in features list or make features clickable to relevant docs.
* Add some screenshots
I spent way too much time clicking every heading to figure out what is this all about till I found Docs link.