Learning Rate Curriculum

Croitoru, Florinel-Alin; Ristea, Nicolae-Catalin; Ionescu, Radu Tudor; Sebe, Nicu

Computer Science > Machine Learning

arXiv:2205.09180 (cs)

[Submitted on 18 May 2022 (v1), last revised 20 Jul 2024 (this version, v4)]

Title:Learning Rate Curriculum

Authors:Florinel-Alin Croitoru, Nicolae-Catalin Ristea, Radu Tudor Ionescu, Nicu Sebe

View PDF HTML (experimental)

Abstract:Most curriculum learning methods require an approach to sort the data samples by difficulty, which is often cumbersome to perform. In this work, we propose a novel curriculum learning approach termed Learning Rate Curriculum (LeRaC), which leverages the use of a different learning rate for each layer of a neural network to create a data-agnostic curriculum during the initial training epochs. More specifically, LeRaC assigns higher learning rates to neural layers closer to the input, gradually decreasing the learning rates as the layers are placed farther away from the input. The learning rates increase at various paces during the first training iterations, until they all reach the same value. From this point on, the neural model is trained as usual. This creates a model-level curriculum learning strategy that does not require sorting the examples by difficulty and is compatible with any neural network, generating higher performance levels regardless of the architecture. We conduct comprehensive experiments on 12 data sets from the computer vision (CIFAR-10, CIFAR-100, Tiny ImageNet, ImageNet-200, Food-101, UTKFace, PASCAL VOC), language (BoolQ, QNLI, RTE) and audio (ESC-50, CREMA-D) domains, considering various convolutional (ResNet-18, Wide-ResNet-50, DenseNet-121, YOLOv5), recurrent (LSTM) and transformer (CvT, BERT, SepTr) architectures. We compare our approach with the conventional training regime, as well as with Curriculum by Smoothing (CBS), a state-of-the-art data-agnostic curriculum learning approach. Unlike CBS, our performance improvements over the standard training regime are consistent across all data sets and models. Furthermore, we significantly surpass CBS in terms of training time (there is no additional cost over the standard training regime for LeRaC). Our code is freely available at: this https URL.

Comments:	Accepted at the International Journal of Computer Vision
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2205.09180 [cs.LG]
	(or arXiv:2205.09180v4 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2205.09180

Submission history

From: Radu Tudor Ionescu [view email]
[v1] Wed, 18 May 2022 18:57:36 UTC (135 KB)
[v2] Sat, 19 Nov 2022 10:41:32 UTC (5,945 KB)
[v3] Fri, 5 Jul 2024 08:51:16 UTC (20,859 KB)
[v4] Sat, 20 Jul 2024 07:29:22 UTC (20,849 KB)

Computer Science > Machine Learning

Title:Learning Rate Curriculum

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Learning Rate Curriculum

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators