Systematic comparison of semi-supervised and self-supervised learning for medical image classification

Huang, Zhe; Jiang, Ruijie; Aeron, Shuchin; Hughes, Michael C.

Computer Science > Computer Vision and Pattern Recognition

arXiv:2307.08919 (cs)

[Submitted on 18 Jul 2023 (v1), last revised 29 Mar 2024 (this version, v3)]

Title:Systematic comparison of semi-supervised and self-supervised learning for medical image classification

Authors:Zhe Huang, Ruijie Jiang, Shuchin Aeron, Michael C. Hughes

View PDF HTML (experimental)

Abstract:In typical medical image classification problems, labeled data is scarce while unlabeled data is more available. Semi-supervised learning and self-supervised learning are two different research directions that can improve accuracy by learning from extra unlabeled data. Recent methods from both directions have reported significant gains on traditional benchmarks. Yet past benchmarks do not focus on medical tasks and rarely compare self- and semi- methods together on an equal footing. Furthermore, past benchmarks often handle hyperparameter tuning suboptimally. First, they may not tune hyperparameters at all, leading to underfitting. Second, when tuning does occur, it often unrealistically uses a labeled validation set that is much larger than the training set. Therefore currently published rankings might not always corroborate with their practical utility This study contributes a systematic evaluation of self- and semi- methods with a unified experimental protocol intended to guide a practitioner with scarce overall labeled data and a limited compute budget. We answer two key questions: Can hyperparameter tuning be effective with realistic-sized validation sets? If so, when all methods are tuned well, which self- or semi-supervised methods achieve the best accuracy? Our study compares 13 representative semi- and self-supervised methods to strong labeled-set-only baselines on 4 medical datasets. From 20000+ GPU hours of computation, we provide valuable best practices to resource-constrained practitioners: hyperparameter tuning is effective, and the semi-supervised method known as MixMatch delivers the most reliable gains across 4 datasets.

Comments:	CVPR 2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:2307.08919 [cs.CV]
	(or arXiv:2307.08919v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2307.08919

Submission history

From: Zhe Huang [view email]
[v1] Tue, 18 Jul 2023 01:31:47 UTC (1,179 KB)
[v2] Sun, 7 Jan 2024 19:07:33 UTC (2,270 KB)
[v3] Fri, 29 Mar 2024 18:19:36 UTC (1,071 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Systematic comparison of semi-supervised and self-supervised learning for medical image classification

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Systematic comparison of semi-supervised and self-supervised learning for medical image classification

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators