Multimodal Dataset Distillation for Image-Text Retrieval

Wu, Xindi; Deng, Zhiwei; Russakovsky, Olga

Computer Science > Computer Vision and Pattern Recognition

arXiv:2308.07545v1 (cs)

[Submitted on 15 Aug 2023 (this version), latest version 7 Feb 2024 (v3)]

Title:Multimodal Dataset Distillation for Image-Text Retrieval

Authors:Xindi Wu, Zhiwei Deng, Olga Russakovsky

View PDF

Abstract:Dataset distillation methods offer the promise of reducing a large-scale dataset down to a significantly smaller set of (potentially synthetic) training examples, which preserve sufficient information for training a new model from scratch. So far dataset distillation methods have been developed for image classification. However, with the rise in capabilities of vision-language models, and especially given the scale of datasets necessary to train these models, the time is ripe to expand dataset distillation methods beyond image classification. In this work, we take the first steps towards this goal by expanding on the idea of trajectory matching to create a distillation method for vision-language datasets. The key challenge is that vision-language datasets do not have a set of discrete classes. To overcome this, our proposed multimodal dataset distillation method jointly distill the images and their corresponding language descriptions in a contrastive formulation. Since there are no existing baselines, we compare our approach to three coreset selection methods (strategic subsampling of the training dataset), which we adapt to the vision-language setting. We demonstrate significant improvements on the challenging Flickr30K and COCO retrieval benchmark: the best coreset selection method which selects 1000 image-text pairs for training is able to achieve only 5.6% image-to-text retrieval accuracy (recall@1); in contrast, our dataset distillation approach almost doubles that with just 100 (an order of magnitude fewer) training pairs.

Comments:	28 pages, 11 figures
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2308.07545 [cs.CV]
	(or arXiv:2308.07545v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2308.07545

Submission history

From: Xindi Wu [view email]
[v1] Tue, 15 Aug 2023 03:22:40 UTC (25,051 KB)
[v2] Mon, 2 Oct 2023 17:50:11 UTC (26,172 KB)
[v3] Wed, 7 Feb 2024 18:57:27 UTC (27,454 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Multimodal Dataset Distillation for Image-Text Retrieval

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Multimodal Dataset Distillation for Image-Text Retrieval

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators