Released Nov 16, 2023Knowledge cutoff Jun 30, 20232,048 context
LLaVA is a large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding, achieving impressive chat capabilities and setting a new state-of-the-art accuracy on Science QA.