MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

← All research

Multimodal Learning & Generation

ICLR 2024 (Poster) · PDF · arXiv · Code · Poster

Let a vision-language model learn from examples containing multiple images. MMICL addresses prompts where images and text are interleaved, where a sentence refers to a particular image, and where the answer depends on relationships between images.

Its context format declares images and assigns image proxy tokens so that textual references and visual representations stay connected. The MIC dataset supplies linked-image tasks and multimodal in-context examples. After image-text pretraining, in-context tuning trains the projection layer and the language model’s query and value parameters while keeping the visual encoder fixed.

MMICL architecture, image proxy tokens and its two-stage training with parameter-freezing legend.
Figure 5: MMICL connects image references to visual embeddings and learns multimodal in-context behavior in two training stages. Click to enlarge.

The paper constructs 5.8 million MIC examples and uses approximately 10% for the reported fine-tuning experiments. Evaluation on MME, MMBench and other vision-language tasks examines multi-image understanding, in-context learning and language bias. The reported gains are relative to the evaluated BLIP-2/InstructBLIP configurations and contemporary baselines.

Resources

Datasets: MIC full · MIC sampled

Checkpoints: T5 XXL · T5 XL

Poster

Download poster (PNG)

MMICL conference poster
Official ICLR conference poster. Click to view the original image.