DiffCap: Exploring Continuous Diffusion on Image Captioning

← All research

Multimodal Learning & Generation

Preprint 2023 · PDF · Code

Generate captions by denoising a whole sentence. DiffCap explores continuous diffusion for image captioning: caption tokens become continuous embeddings, noise is added during training, and a denoising model recovers the sentence conditioned on visual features.

A frozen visual encoder provides the image representation. A projected image feature conditions a BERT-style denoiser together with a diffusion timestep embedding, while an output head converts the recovered embeddings back to discrete tokens. The caption’s text embeddings are diffused; the image condition stays fixed.

DiffCap architecture showing a fixed visual condition and denoising of caption token embeddings.
Figure 1: Image-conditioned continuous diffusion over caption embeddings, with denoising and token prediction losses. Click to enlarge.

On the COCO Karpathy test split, the reported model reaches 31.6 BLEU-4 and 104.3 CIDEr, with caption diversity evaluated separately using inter-distinct and self-BLEU. The study finds a quality/diversity trade-off: captioning quality is comparable to the tested non-autoregressive baselines, while the strongest evaluated autoregressive systems still score higher on standard quality metrics.