I'd be excited to see what a more substantial, state-of-the-art model could do with geospatial data.
Say, a model based on something like ViT-22B, with 22 billion parameters: https://arxiv.org/abs/2302.05442.
I'd be excited to see what a more substantial, state-of-the-art model could do with geospatial data.
Say, a model based on something like ViT-22B, with 22 billion parameters: https://arxiv.org/abs/2302.05442.
With LLMs taking over the spotlight it's easy for people to forget that not everything needs billions or trillions of parameters. Stable Diffusion fits comfortably in my 8GB of VRAM and can generate amazing images. I'd love to see more research like this in smaller models that can be used on cheap consumer hardware.
If a model is going to be used many times for a specific use case, it is far cheaper and uses far less energy to fine tune a small model once and run it on cheap low-power hardware than it is to continuously run a huge, do-everything model on expensive, high-power hardware. Enormous models are great for exploration and for general purpose applications like ChatGPT, but I think that we will find over the next few years that smaller, purpose-built models will continue to dominate in applications like geospatial analysis.
As I understand it we're contrasting two opposite approaches to ML: fine tuning small models for specific applications versus training a single large model that can generalize to new tasks without preparing them ahead of time.
I'm arguing that in fine tuning is far more useful than people are currently giving it credit for, and that generalizing a single massive model to new tasks is overrated.
Can you clarify where you're seeing a straw man?
I'm not arguing that there is no place for large, general models—they're great for exploration—just that a smaller foundation model shouldn't be dismissed offhand based solely on parameter count.
I think its just that there is a dearth of publishing on the advances in geospatial ml.