Iterate averaging as regularization for stochastic gradient descent

Abstract

We propose and analyze a variant of the classic Polyak-Ruppert averaging scheme, broadly used in stochastic gradient methods. Rather than a uniform average of the iterates, we consider a weighted average, with weights decaying in a geometric fashion. In the context of linear least squares regression, we show that this averaging scheme has a the same regularizing effect, and indeed is asymptotically equivalent, to ridge regression. In particular, we derive finite-sample bounds for the proposed approach that match the best known results for regularized stochastic gradient methods.

BibTeX key: neu2018iterate
entry type: article
year: 2018
url: http://arxiv.org/abs/1802.08009
note: cite arxiv:1802.08009

Users

Comments and Reviewsshow / hide

Please log in to take part in the discussion (add own reviews or comments).

BibSonomy

Iterate averaging as regularization for stochastic gradient descent

Abstract

Tags

Users

Comments and Reviewsshow / hide

Cite this publication

More citation styles

search on