LLaVA is a large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding, achieving impressive chat capabilities and setting a new state-of-the-art accuracy on Science QA.
#multimodal
Modalities
Context
2K
Released
Nov 16, 2023
Knowledge Cutoff
Jun 2023
Token volume and request traffic to this model over time.
LLaVA is a large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding, achieving impressive chat capabilities and setting a new state-of-the-art accuracy on Science QA. #multimodal
LLaVA 13B has a 2,048 token context window.
LLaVA 13B accepts text and images as input and returns text.
LLaVA 13B was released on November 16, 2023. Its knowledge cutoff is June 30, 2023.