Wow! You folks are making huge strides for open-source multimodal models. Thank you for all the time and effort on these as they will open up many opportunities for researchers and developers. Also, the emergent zero-shot capabilities when LLaVA-1.6 is tested against Chinese benchmarks with only English multi-modal training data are interesting and that may be a good direction for future research.