DREAM: Efficient Dataset Distillation by Representative Matching

Liu, Yanqing; Gu, Jianyang; Wang, Kai; Zhu, Zheng; Jiang, Wei; You, Yang

Computer Science > Computer Vision and Pattern Recognition

arXiv:2302.14416v1 (cs)

[Submitted on 28 Feb 2023 (this version), latest version 30 Aug 2023 (v3)]

Title:DREAM: Efficient Dataset Distillation by Representative Matching

Authors:Yanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu, Wei Jiang, Yang You

View PDF

Abstract:Dataset distillation aims to generate small datasets with little information loss as large-scale datasets for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample generation process by matching synthetic images and the original ones regarding gradients, embedding distributions, or training trajectories. Although there are various matching objectives, currently the method for selecting original images is limited to naive random sampling. We argue that random sampling inevitably involves samples near the decision boundaries, which may provide large or noisy matching targets. Besides, random sampling cannot guarantee the evenness and diversity of the sample distribution. These factors together lead to large optimization oscillations and degrade the matching efficiency. Accordingly, we propose a novel matching strategy named as \textbf{D}ataset distillation by \textbf{RE}present\textbf{A}tive \textbf{M}atching (DREAM), where only representative original images are selected for matching. DREAM is able to be easily plugged into popular dataset distillation frameworks and reduce the matching iterations by 10 times without performance drop. Given sufficient training time, DREAM further provides significant improvements and achieves state-of-the-art performances.

Comments:	Efficient matching for dataset distillation
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2302.14416 [cs.CV]
	(or arXiv:2302.14416v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2302.14416

Submission history

From: Kai Wang [view email]
[v1] Tue, 28 Feb 2023 08:48:45 UTC (4,125 KB)
[v2] Thu, 9 Mar 2023 15:53:56 UTC (3,670 KB)
[v3] Wed, 30 Aug 2023 14:22:32 UTC (3,668 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:DREAM: Efficient Dataset Distillation by Representative Matching

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:DREAM: Efficient Dataset Distillation by Representative Matching

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators