{
  "id": 20201,
  "title": "Why no regularization on convolutional layers?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20201",
  "author_name": "",
  "post_date": "2016-04-17T16:38:40.450Z",
  "votes": 4,
  "comment_count": 4,
  "views": 2092,
  "content": "<p>I'm fairly new to convolutional neural networks and have a question after looking at some of the scripts that have been posted. (Many thanks to those posting them!)</p>\n\n<p>Those scripts typically add dropout regularization to the fully connected layers, but I don't see any regularization on the convolutional layers. When I read about CNNs (such as in the <a href=\"http://arxiv.org/pdf/1409.1556.pdf\">VGG model paper</a>) they use L2 regularization on those layers.</p>\n\n<p>Is the omission of L2 regularization just a simplification in the example scripts or am I missing something?</p>",
  "messages": [
    {
      "id": "115275",
      "postDate": "04/17/2016 16:38:40",
      "content": "<p>I'm fairly new to convolutional neural networks and have a question after looking at some of the scripts that have been posted. (Many thanks to those posting them!)</p>\n\n<p>Those scripts typically add dropout regularization to the fully connected layers, but I don't see any regularization on the convolutional layers. When I read about CNNs (such as in the <a href=\"http://arxiv.org/pdf/1409.1556.pdf\">VGG model paper</a>) they use L2 regularization on those layers.</p>\n\n<p>Is the omission of L2 regularization just a simplification in the example scripts or am I missing something?</p>",
      "rawMarkdown": "I'm fairly new to convolutional neural networks and have a question after looking at some of the scripts that have been posted. (Many thanks to those posting them!)\r\n\r\nThose scripts typically add dropout regularization to the fully connected layers, but I don't see any regularization on the convolutional layers. When I read about CNNs (such as in the [VGG model paper][1]) they use L2 regularization on those layers.\r\n\r\nIs the omission of L2 regularization just a simplification in the example scripts or am I missing something?\r\n\r\n\r\n  [1]: http://arxiv.org/pdf/1409.1556.pdf",
      "votes": null
    },
    {
      "id": "115417",
      "postDate": "04/18/2016 16:18:08",
      "content": "<p>You can use other regularizers in addition to dropout but dropout works well enough on its own and L1/L2 are more difficult to tune (compared to dropout). Additionally, adjustments to the architecture may require adjusting the L1/L2 regularization. Towards the end of the competition, it may be useful to apply and tune other regularization methods.</p>\n\n<p>Applying dropout to the final fully-connected layers effectively ensemble the entire network, including all previous layers.</p>\n\n<p>Dropout has the added benefit of reducing dependencies within each layer so it can be beneficial to apply to all of them. For example, in the original paper it was applied to convolutional layers in addition to fully-connected layers and showed improvement over just applying dropout to the fully-connected layers. However, dropout can increase training time so it is usually omitted from the convolutional layers.</p>\n\n<p>Furthermore, batch normalization largely removes the need for dropout (see bn paper) but, due to the size of the dataset here, keeping dropout on the fully-connected layers (which don't use bn) was helpful.</p>\n\n<p>[0] <a href=\"http://www.cs.toronto.edu/~rsalakhu/papers/srivastava14a.pdf\">http://www.cs.toronto.edu/~rsalakhu/papers/srivastava14a.pdf</a></p>",
      "rawMarkdown": "You can use other regularizers in addition to dropout but dropout works well enough on its own and L1/L2 are more difficult to tune (compared to dropout). Additionally, adjustments to the architecture may require adjusting the L1/L2 regularization. Towards the end of the competition, it may be useful to apply and tune other regularization methods.\r\n\r\nApplying dropout to the final fully-connected layers effectively ensemble the entire network, including all previous layers.\r\n\r\nDropout has the added benefit of reducing dependencies within each layer so it can be beneficial to apply to all of them. For example, in the original paper it was applied to convolutional layers in addition to fully-connected layers and showed improvement over just applying dropout to the fully-connected layers. However, dropout can increase training time so it is usually omitted from the convolutional layers.\r\n\r\nFurthermore, batch normalization largely removes the need for dropout (see bn paper) but, due to the size of the dataset here, keeping dropout on the fully-connected layers (which don't use bn) was helpful.\r\n\r\n[0] http://www.cs.toronto.edu/~rsalakhu/papers/srivastava14a.pdf",
      "votes": null
    },
    {
      "id": "115424",
      "postDate": "04/18/2016 16:49:21",
      "content": "<p>Deleted my own comment, it was not useful - I think I just confused myself with something else I read (would be nice to be able to hide posts made in error!)</p>",
      "rawMarkdown": "Deleted my own comment, it was not useful - I think I just confused myself with something else I read (would be nice to be able to hide posts made in error!)",
      "votes": null
    },
    {
      "id": "115456",
      "postDate": "04/18/2016 18:46:16",
      "content": "<p>@Jim Fleming: Thanks for that very clear explanation. </p>\n\n<p>I experimented with adding L2 regularization on the convolutional layers and it didn't help much, if at all, confirming what you said.</p>",
      "rawMarkdown": "Jim Fleming: Thanks for that very clear explanation. \r\n\r\nI experimented with adding L2 regularization on the convolutional layers and it didn't help much, if at all, confirming what you said.",
      "votes": null
    },
    {
      "id": "157159",
      "postDate": "01/19/2017 14:01:24",
      "content": "<p>To my understanding it seems controversial to use regularization in the conv layers. We use regularization to prevent overfitting. So that a network doesn't memorize the training set and can't do well when predicting outside of the train data. But in conv layer we don't do any prediction. We actually extract features with these layers. So if we add regularization like dropout we are actually hampering the process of extracting those features effectively. There is no overfitting issue with this layer as it doesn't do prediction so there should be no need for regularization on these layers. </p>\n\n<p>This is my assumption but I'm new in this field and have lot to learn. If anyone can point out if I'm mistaken or can elaborate my points more accurately it would be greatly appreciated.</p>",
      "rawMarkdown": "To my understanding it seems controversial to use regularization in the conv layers. We use regularization to prevent overfitting. So that a network doesn't memorize the training set and can't do well when predicting outside of the train data. But in conv layer we don't do any prediction. We actually extract features with these layers. So if we add regularization like dropout we are actually hampering the process of extracting those features effectively. There is no overfitting issue with this layer as it doesn't do prediction so there should be no need for regularization on these layers. \r\n\r\nThis is my assumption but I'm new in this field and have lot to learn. If anyone can point out if I'm mistaken or can elaborate my points more accurately it would be greatly appreciated.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115417,
      "author_name": "jimfleming",
      "author_url": "",
      "post_date": "04/18/2016 16:18:08",
      "content": "<p>You can use other regularizers in addition to dropout but dropout works well enough on its own and L1/L2 are more difficult to tune (compared to dropout). Additionally, adjustments to the architecture may require adjusting the L1/L2 regularization. Towards the end of the competition, it may be useful to apply and tune other regularization methods.</p>\n\n<p>Applying dropout to the final fully-connected layers effectively ensemble the entire network, including all previous layers.</p>\n\n<p>Dropout has the added benefit of reducing dependencies within each layer so it can be beneficial to apply to all of them. For example, in the original paper it was applied to convolutional layers in addition to fully-connected layers and showed improvement over just applying dropout to the fully-connected layers. However, dropout can increase training time so it is usually omitted from the convolutional layers.</p>\n\n<p>Furthermore, batch normalization largely removes the need for dropout (see bn paper) but, due to the size of the dataset here, keeping dropout on the fully-connected layers (which don't use bn) was helpful.</p>\n\n<p>[0] <a href=\"http://www.cs.toronto.edu/~rsalakhu/papers/srivastava14a.pdf\">http://www.cs.toronto.edu/~rsalakhu/papers/srivastava14a.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115424,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/18/2016 16:49:21",
      "content": "<p>Deleted my own comment, it was not useful - I think I just confused myself with something else I read (would be nice to be able to hide posts made in error!)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115456,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "04/18/2016 18:46:16",
      "content": "<p>@Jim Fleming: Thanks for that very clear explanation. </p>\n\n<p>I experimented with adding L2 regularization on the convolutional layers and it didn't help much, if at all, confirming what you said.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157159,
      "author_name": "calicratis19",
      "author_url": "",
      "post_date": "01/19/2017 14:01:24",
      "content": "<p>To my understanding it seems controversial to use regularization in the conv layers. We use regularization to prevent overfitting. So that a network doesn't memorize the training set and can't do well when predicting outside of the train data. But in conv layer we don't do any prediction. We actually extract features with these layers. So if we add regularization like dropout we are actually hampering the process of extracting those features effectively. There is no overfitting issue with this layer as it doesn't do prediction so there should be no need for regularization on these layers. </p>\n\n<p>This is my assumption but I'm new in this field and have lot to learn. If anyone can point out if I'm mistaken or can elaborate my points more accurately it would be greatly appreciated.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115275": "I'm fairly new to convolutional neural networks and have a question after looking at some of the scripts that have been posted. (Many thanks to those posting them!)\r\n\r\nThose scripts typically add dropout regularization to the fully connected layers, but I don't see any regularization on the convolutional layers. When I read about CNNs (such as in the [VGG model paper][1]) they use L2 regularization on those layers.\r\n\r\nIs the omission of L2 regularization just a simplification in the example scripts or am I missing something?\r\n\r\n\r\n  [1]: http://arxiv.org/pdf/1409.1556.pdf",
    "115417": "You can use other regularizers in addition to dropout but dropout works well enough on its own and L1/L2 are more difficult to tune (compared to dropout). Additionally, adjustments to the architecture may require adjusting the L1/L2 regularization. Towards the end of the competition, it may be useful to apply and tune other regularization methods.\r\n\r\nApplying dropout to the final fully-connected layers effectively ensemble the entire network, including all previous layers.\r\n\r\nDropout has the added benefit of reducing dependencies within each layer so it can be beneficial to apply to all of them. For example, in the original paper it was applied to convolutional layers in addition to fully-connected layers and showed improvement over just applying dropout to the fully-connected layers. However, dropout can increase training time so it is usually omitted from the convolutional layers.\r\n\r\nFurthermore, batch normalization largely removes the need for dropout (see bn paper) but, due to the size of the dataset here, keeping dropout on the fully-connected layers (which don't use bn) was helpful.\r\n\r\n[0] http://www.cs.toronto.edu/~rsalakhu/papers/srivastava14a.pdf",
    "115424": "Deleted my own comment, it was not useful - I think I just confused myself with something else I read (would be nice to be able to hide posts made in error!)",
    "115456": "Jim Fleming: Thanks for that very clear explanation. \r\n\r\nI experimented with adding L2 regularization on the convolutional layers and it didn't help much, if at all, confirming what you said.",
    "157159": "To my understanding it seems controversial to use regularization in the conv layers. We use regularization to prevent overfitting. So that a network doesn't memorize the training set and can't do well when predicting outside of the train data. But in conv layer we don't do any prediction. We actually extract features with these layers. So if we add regularization like dropout we are actually hampering the process of extracting those features effectively. There is no overfitting issue with this layer as it doesn't do prediction so there should be no need for regularization on these layers. \r\n\r\nThis is my assumption but I'm new in this field and have lot to learn. If anyone can point out if I'm mistaken or can elaborate my points more accurately it would be greatly appreciated."
  },
  "source": "meta"
}