TENGraD: Time-Efficient Natural Gradient Descent with Exact Fisher-Block Inversion

Soori, Saeed; Can, Bugra; Mu, Baourun; Gürbüzbalaban, Mert; Dehnavi, Maryam Mehri

Computer Science > Machine Learning

arXiv:2106.03947 (cs)

[Submitted on 7 Jun 2021 (v1), last revised 3 Mar 2022 (this version, v4)]

Title:TENGraD: Time-Efficient Natural Gradient Descent with Exact Fisher-Block Inversion

Authors:Saeed Soori, Bugra Can, Baourun Mu, Mert Gürbüzbalaban, Maryam Mehri Dehnavi

View PDF

Abstract:This work proposes a time-efficient Natural Gradient Descent method, called TENGraD, with linear convergence guarantees. Computing the inverse of the neural network's Fisher information matrix is expensive in NGD because the Fisher matrix is large. Approximate NGD methods such as KFAC attempt to improve NGD's running time and practical application by reducing the Fisher matrix inversion cost with approximation. However, the approximations do not reduce the overall time significantly and lead to less accurate parameter updates and loss of curvature information. TENGraD improves the time efficiency of NGD by computing Fisher block inverses with a computationally efficient covariance factorization and reuse method. It computes the inverse of each block exactly using the Woodbury matrix identity to preserve curvature information while admitting (linear) fast convergence rates. Our experiments on image classification tasks for state-of-the-art deep neural architecture on CIFAR-10, CIFAR-100, and Fashion-MNIST show that TENGraD significantly outperforms state-of-the-art NGD methods and often stochastic gradient descent in wall-clock time.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2106.03947 [cs.LG]
	(or arXiv:2106.03947v4 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2106.03947

Submission history

From: Saeed Soori [view email]
[v1] Mon, 7 Jun 2021 20:16:15 UTC (34,813 KB)
[v2] Fri, 27 Aug 2021 17:42:03 UTC (1 KB) (withdrawn)
[v3] Wed, 23 Feb 2022 20:05:41 UTC (17,401 KB)
[v4] Thu, 3 Mar 2022 15:10:43 UTC (17,401 KB)

Computer Science > Machine Learning

Title:TENGraD: Time-Efficient Natural Gradient Descent with Exact Fisher-Block Inversion

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:TENGraD: Time-Efficient Natural Gradient Descent with Exact Fisher-Block Inversion

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators