Presentation Information

[N-1-18]Polyak-Ruppert Averaging No Full Grad SVRG

◎△Suguru Asai1, Ryo Yamatomi1, Hiroshi Ninomiya1 (1. Shonan Institute of Technology)

Keywords:

Stochastic Gradient Descent,Stochastic Variance Reduced Gradient,Polyak-Ruppert Averaging,Gradient Variance Reduction,Deep Learning Optimization

Stochastic Gradient Descent (SGD) is widely used for deep learning optimization, but the stochastic gradients computed from mini-batches often exhibit large variance, resulting in unstable training. Stochastic Variance Reduced Gradient (SVRG) reduces gradient variance by using the full gradient at a reference point. To avoid the high computational cost of full-gradient evaluation, No Full Grad SVRG approximates the full gradient with the average gradient from the previous epoch. However, since it employs the last parameter of the previous epoch as the reference point, a discrepancy between the full gradient and the average gradient may arise, leading to unstable learning. In this paper, we propose Polyak-Ruppert Averaging No Full Grad SVRG (Polyak-SVRG), which uses the averaged parameter of the previous epoch as the reference point. The averaged parameter smooths abrupt fluctuations caused by individual updates and improves the approximation accuracy of the full gradient by the average gradient. Experimental results on the CIFAR-10 image classification task demonstrate that the proposed method achieves more stable learning with smaller oscillations in prediction accuracy than existing methods.