What’s the best library for fine-tuning VLMs at the moment and do they support this architecture or that for the IBM Granite vision models? Document understanding tasks seem in special need of fine-tuning.
It was fine tuned from this: https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct
There's an example of fine tuning the base that would likely be applicable to this one as well.