MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Multimodal Learning & Generation
ICLR 2024 (Poster) · PDF · arXiv · Code · Poster
Let a vision-language model learn from examples containing multiple images. MMICL addresses prompts where images and text are interleaved, where a sentence refers to a particular image, and where the answer depends on relationships between images.
Its context format declares images and assigns image proxy tokens so that textual references and visual representations stay connected. The MIC dataset supplies linked-image tasks and multimodal in-context examples. After image-text pretraining, in-context tuning trains the projection layer and the language model’s query and value parameters while keeping the visual encoder fixed.
The paper constructs 5.8 million MIC examples and uses approximately 10% for the reported fine-tuning experiments. Evaluation on MME, MMBench and other vision-language tasks examines multi-image understanding, in-context learning and language bias. The reported gains are relative to the evaluated BLIP-2/InstructBLIP configurations and contemporary baselines.
Resources
Datasets: MIC full · MIC sampled
Poster
