DiffCap: Exploring Continuous Diffusion on Image Captioning
Multimodal Learning & Generation
Generate captions by denoising a whole sentence. DiffCap explores continuous diffusion for image captioning: caption tokens become continuous embeddings, noise is added during training, and a denoising model recovers the sentence conditioned on visual features.
A frozen visual encoder provides the image representation. A projected image feature conditions a BERT-style denoiser together with a diffusion timestep embedding, while an output head converts the recovered embeddings back to discrete tokens. The caption’s text embeddings are diffused; the image condition stays fixed.
On the COCO Karpathy test split, the reported model reaches 31.6 BLEU-4 and 104.3 CIDEr, with caption diversity evaluated separately using inter-distinct and self-BLEU. The study finds a quality/diversity trade-off: captioning quality is comparable to the tested non-autoregressive baselines, while the strongest evaluated autoregressive systems still score higher on standard quality metrics.
