IBM and NASA Open Source Largest Geospatial AI Foundation Model on Hugging Face
newsroom.ibm.com
newsroom.ibm.com
However, make no mistake: this is for the scientific community and will not help geospatial data to be commercialized. No one cares about your geospatial crop model or that you can identify energy infrastructure or that there's some activity around that copper mine. Well, at least no one cares that will actually pay you.
(FWIW, I cofounded a geospatial analytics company)
Satellite data is extremely idiosyncratic. It's coarse (~10m at best), infrequent (every few days at best), and oh you have to deal with the fact that the planet is covered in 50% clouds at any moment. Satellite data works best on things that don't move, that are fairly large, and change infrequently. If you find a use case that satisfies those conditions and want to make money, then you need to find a problem that terrestrial sensors haven't solved. And if you find that problem, the cost of building, training, and running your model (plus the cost of the data!) has to be less than the marginal value of your model. Good luck finding those use cases.
The US Government is special. We don't know what's going on in North Korea or Ukraine or the South China Sea so we buy high resolution imagery over those areas (30cm) at great cost. Large ag companies and oil companies know what's going on within their own facilities; and price gives them information about the rest of the supply chain.
In other words, this might be an interesting announcement for scientists, but it won't change the geospatial market at all.
I'm not far from bashing the VC scene and the adjacent startup culture, but your overly cynical comment was too much even for me. More intellectual humility and less cheap soundbites would benefit society a lot.
If you're interested, he wrote about it: https://philosophygeek.medium.com/meditations-a-requiem-for-...
Oh man. I had no idea! These guys were some of my prime competition for years.
The true cost of venture capitol revealed.
Are the coarseness and cloud aspects going to become less of a factor now that there are commercial high-resolution synthetic aperture radar imagery providers? I'm just a hobbyist, but the imagery I've seen is <i>sharp</i>, and it even caught the NRO's attention.[1]
[1] https://spacenews.com/national-reconnaissance-office-signs-a...
InfSar (InSar, SAR, whatever we're calling it these days) isn't a drop in replacement for anything. Its really neither here nor there when it comes to the utility of other dataset. Infsar is amazing, dgmw, but its stands on its own and has its own advantages/ disadvantages.
The ocs point stands. Satellite data is tough because there is a shit ton of atmosphere between you and the target. That issue doesn't go away with infsar and especially not if it isnt coincidentally collected with higher resolution spectral data. I've been in the industry for around 15 years. Things have gotten better, but really, its important to understand the context and limitations of specific platforms. Afaik, there is no panacea.
>The best commercially available spatial resolution for optical imagery is 25 cm, which means that one pixel represents a 25-by-25-cm area on the ground—roughly the size of your laptop.
There are many very nice commercial products available.
Disclaimer: I work for Planet.
I also disagree with the assertion "no one will actually pay you." Read pages 26-28 of the quarterly report for more information.
NYSE: PL
That being said, while I do believe that Planet has one of the best business models in the industry, I do sometimes worry that they are a bit early, and their customers aren't ready for them yet.
As someone who has to work with their data (among others), Planet has some of the best APIs in the industry (it's a low bar though).
I'm mostly just playing devils advocate here, and the point I'm trying to make is share price alone isn't everything. Planet is growing its revenue YoY, and they might even be profitable soon (lol).
Also, my original point was that the company I work for buys a lot of data from Planet, which is something that's just increasing as we grow.
> Satellite data is extremely idiosyncratic. It's coarse (~10m at best) 10m is the best free imagery, in the commercial domain it goes down to 30cm.
> Well, at least no one cares that will actually pay you. There are plenty of things people will pay you for. But you gotta find those niches.
> In other words, this might be an interesting announcement for scientists, but it won't change the geospatial market at all.
Maybe, we'll definitely check if it can be fine-tuned on higher res data. We do sometimes use Sentinel-2 (not a lot though), but can help with those cases.
- https://huggingface.co/spaces/ibm-nasa-geospatial/Prithvi-10... - This demo showcases how the model was finetuned to detect water at a higher resolution than it was trained on (i.e. 10m versus 30m) using Sentinel 2 imagery from on the sen1floods11 dataset
- https://huggingface.co/spaces/ibm-nasa-geospatial/Prithvi-10... - This demo showcases how the model was finetuned to classify crop and other land use categories using multi temporal data.
- https://huggingface.co/spaces/ibm-nasa-geospatial/Prithvi-10... - This demo showcases the image reconstracting over three timestamps, with the user providing a set of three HLS images and the model randomly masking out some proportion of the images and then reconstructing them based on the not masked portion of the images
- https://huggingface.co/spaces/ibm-nasa-geospatial/Prithvi-10... - This demo showcases how the model was finetuned to detect burn scars
More/same but different source information:
- From NASA - https://www.earthdata.nasa.gov/news/impact-ibm-hls-foundatio...
- From IBM - https://research.ibm.com/blog/nasa-hugging-face-ibm
- From huggingface - https://huggingface.co/ibm-nasa-geospatial
I prefer the full press release, especially since it already has the link to Hugging Face for those who want it.
> The model – trained jointly by IBM and NASA on Harmonized Landsat Sentinel-2 satellite data (HLS) over one year across the continental United States and fine-tuned on labeled data for flood and burn scar mapping — has demonstrated to date a 15 percent improvement over state-of-the-art techniques using half as much labeled data. With additional fine tuning, the base model can be redeployed for tasks like tracking deforestation, predicting crop yields, or detecting and monitoring greenhouse gasses. IBM and NASA researchers are also working with Clark University to adapt the model for applications such as time-series segmentation and similarity research.
There are larger foundation models for geospatial imagery available. Our pre-training method, Scale-MAE [0], has 323M parameters, makes encoders robust to changes in satellite imagery resolution, and is therefore trained on satellite imagery of all resolutions. Work out of SI Analytics [1] presents a 2.4B parameter transformers for satellite imagery.
The IBM offering is appealing not because it is "first" or the "best", but because it is accessible. A lot of institutions / enterprises (including my own) are able to leverage transformer models because HF has done so much to lower the barrier to entry with their hub, documentation (model cards) and API.
How else would I have found your models but for an IBM press release?
For example: you can easily use Photoshop to distinguish between various hues, apply a mask based on the selection, and then apply a color overlay, essentially "highlighting" a particular hue with another. (I imagine a similar approach can be used with hyperspectral imaging as well.)
As visual data is color-based, why can't such "dumb" filters be used? Why is this an AI challenge, let alone one well suited to the flexibility of foundation models?
*Edit: I believe I may have answered my own question. I assume the advantage this approach has is in the efficient bulk analysis of geospatial data. Platform consumes visual data, and spits out numerical data based on an analysis of the images. That numerical data can then be manipulated and fed into prediction models.
Enter remotely sensed dataset. They are monumental in scale. So much so that the 804 chips they submitted here is pretty much laughable. Likewise, some parameters of the data have inherent meaning; the pixel dimensions correspond to real world measurements; the bands are specific spectral windows.
Its similar but far more than just working with RGB cell phone camera data.
I think partnering with Huggingface was a good move, because it means the interface is easy to use. This difficulty in actually using pre-trained research models was one of the design goals of Moonshine[1] and I have no doubt that if it was IBM alone it wouldn't be nearly as easy to use.
Will be excited to hear if this works for people! Always cool to see your idea validated even if nobody really uses your tool :)
With the better resolutions that are being launched and current AI, there are many more feasible applications vs. when you started DL.
We've built this leveraging other foundation models, so all research is very much appreciated https://www.youtube.com/watch?v=2yz4DwPtdjE
Disclaimer: I'm a co-founder of Happyrobot. We're working on this as we speak.
Let's all meet at SmallSat https://smallsat.org/ conference next week and discuss in more depth!
But, could this "model" be used for something like monitoring land use in a city? The specific example I'm thinking of is getting a percentage breakdown of what land devoted to paved surfaces (parking/roads), to vacant undeveloped lots, and to built structures. It would also be interesting to see how those percentages have changed over time.
That may not be fine enough resolution in the source data to resolve parking lots Vs undeveloped areas to the degree you require.
The alternatives are to use the nature of the models (trained to lock in on multispectral signatures) on finer source data - which may limit your options about the globe, OR
to use urban data from city land agencies which have maps from low level air photo surveys with resolution down to 10 cm (+/-) and GIS land boundaries which are often classified via metadata.
Air photos are not multispectral (usually) so you won't have access to IR bands etc.
You can get coarse city growth figures globally from sat data going back to TERRA (launched 1999 IIRC) and fine grained air photo data from well off cities going back to (say) the mid 80s (and longer for wet negatives).
The typical trick is to look for areas which absorb visible red while reflect near infrared to identify vegetation. If all you have is rgb imagery then you can use machine learning techniques to develop a classification system.
It does not look like this model means a breakthrough in your application area. Definietly not out of the box, maybe with more work you can refine it to do the classification you are looking for.
Do you have access to satelite images of areas where you would be interested in these percentages?
Answering my own question: it seems one can access the right type of data from the Sentinel satelites relatively freely.
I'd be excited to see what a more substantial, state-of-the-art model could do with geospatial data.
Say, a model based on something like ViT-22B, with 22 billion parameters: https://arxiv.org/abs/2302.05442.
With LLMs taking over the spotlight it's easy for people to forget that not everything needs billions or trillions of parameters. Stable Diffusion fits comfortably in my 8GB of VRAM and can generate amazing images. I'd love to see more research like this in smaller models that can be used on cheap consumer hardware.
If a model is going to be used many times for a specific use case, it is far cheaper and uses far less energy to fine tune a small model once and run it on cheap low-power hardware than it is to continuously run a huge, do-everything model on expensive, high-power hardware. Enormous models are great for exploration and for general purpose applications like ChatGPT, but I think that we will find over the next few years that smaller, purpose-built models will continue to dominate in applications like geospatial analysis.
As I understand it we're contrasting two opposite approaches to ML: fine tuning small models for specific applications versus training a single large model that can generalize to new tasks without preparing them ahead of time.
I'm arguing that in fine tuning is far more useful than people are currently giving it credit for, and that generalizing a single massive model to new tasks is overrated.
Can you clarify where you're seeing a straw man?
I'm not arguing that there is no place for large, general models—they're great for exploration—just that a smaller foundation model shouldn't be dismissed offhand based solely on parameter count.
I think its just that there is a dearth of publishing on the advances in geospatial ml.
Floods and fires are bad news, but how does this product actually get used to make things any different ?
E.g. https://huggingface.co/ibm-nasa-geospatial/Prithvi-100M
(Though with all Facebook's shenanigans I understand the need to be sceptical and search for a license when a company claims "open source")
Yes, I actually type these by hand. My carping is 100 percent artisanal.
> With additional fine tuning, the base model can be redeployed for tasks like tracking deforestation, predicting crop yields, or detecting and monitoring greenhouse gasses. IBM and NASA researchers are also working with Clark University to adapt the model for applications such as time-series segmentation and similarity research.
When I worked on ML inference I would tease the researchers with the question “what is a model?” in the hope they would say something that constrained it in any way, but no, a model can be anything at all and is whatever you want it to be.
As such this press release is perfect nonsense, which is kind of appropriate for a Watson offshoot. If there is something interesting here it doesn’t succeed in telling you what it is.
“Prithvi is a first-of-its-kind temporal Vision transformer”
That is enormously more specific and interesting, and the word model is absent.
In contrast, if you’re publishing this in a machine learning journal, then sure—you should definitely call it a temporal Vision transformer.
Or some kind of weird robot-car hybrid?
I guess since it's a "Vision" transformer, it's probably actually some kind of hallucinogenic plant.
It's a real shame that the term is being abused so badly, because it's really appropriate in a lot of these cases. Or rather, it would be if people used it mindfully.
It very literally is a model of the relationships within the training data.
Moreover, statisticians and other varieties of people doing data analysis have been using the term "model" to mean something along these lines for many decades already.
If anything, it's nice to see more people calling these things "models".