Block layer decomposition schemes for training deep neural networks

Academic Article

Publication Date:

2020

abstract:

Deep feedforward neural networks' (DFNNs) weight estimation relies on the solution of a very large nonconvex optimization problem that may have many local (no global) minimizers, saddle points and large plateaus. Furthermore, the time needed to find good solutions of the training problem heavily depends on both the number of samples and the number of weights (variables). In this work, we show how block coordinate descent (BCD) methods can be fruitful applied to DFNN weight optimization problem and embedded in online frameworks possibly avoiding bad stationary points. We first describe a batch BCD method able to effectively tackle difficulties due to the network's depth; then we further extend the algorithm proposing an online BCD scheme able to scale with respect to both the number of variables and the number of samples. We perform extensive numerical results on standard datasets using various deep networks. We show that the application of BCD methods to the training problem of DFNNs improves over standard batch/online algorithms in the training phase guaranteeing good generalization performance as well.

Iris type:

01.01 Articolo in rivista

Keywords:

Deep feedforward neural networks; Block coordinate decomposition; Online optimization; Large scale optimization

List of contributors:

Palagi, Laura

Handle:

https://iris.cnr.it/handle/20.500.14243/382240

Published in:

JOURNAL OF GLOBAL OPTIMIZATION

Journal