{
  "id": 132517,
  "title": "Weight decay in Adam optimizer",
  "url": "/competitions/bengaliai-cv19/discussion/132517",
  "author_name": "Kupchanski",
  "post_date": "2020-02-26T12:23:53.344000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Is it helpful to use WD in Adam optimizer? I heard that it isn`t work correctly. I am using pytorch btw.</p>",
  "messages": [
    {
      "id": 759075,
      "postDate": "2020-02-28T14:31:04.310Z",
      "content": "<p>You should definitely check this paper: Decoupled Weight Decay Regularization ( <a href=\"https://arxiv.org/abs/1711.05101\">https://arxiv.org/abs/1711.05101</a> ).</p>",
      "rawMarkdown": "You should definitely check this paper: Decoupled Weight Decay Regularization ( https://arxiv.org/abs/1711.05101 ).",
      "votes": 1
    },
    {
      "id": 758926,
      "postDate": "2020-02-28T10:56:09.660Z",
      "content": "<p><a href=\"https://youtu.be/_JB0AO7QxSA?list=PLf7L7Kg8_FNxHATtLwDceyh72QQL9pvpQ&amp;t=1910\">here</a> Justin Johnson says that: there is theoritical work that most shallow minimums  that SGD-momentum misses are actually bad minimums - minimums that does not generalize well on test set, that we call local minimum, so because Adam optimizer is based on SGD momentum it also have that \"feature\", that it can jump over bad minimum,. So you can use regularization and don't care about landing in local minimum. Also for other side minimums that  SGD-momentum skips is not always bad minimums and it sometimes can be  best to land there so by regularization parameter you have control on this \"feature\".</p>",
      "rawMarkdown": "[here](https://youtu.be/_JB0AO7QxSA?list=PLf7L7Kg8_FNxHATtLwDceyh72QQL9pvpQ&amp;t=1910) Justin Johnson says that: there is theoritical work that most shallow minimums  that SGD-momentum misses are actually bad minimums - minimums that does not generalize well on test set, that we call local minimum, so because Adam optimizer is based on SGD momentum it also have that \"feature\", that it can jump over bad minimum,. So you can use regularization and don't care about landing in local minimum. Also for other side minimums that  SGD-momentum skips is not always bad minimums and it sometimes can be  best to land there so by regularization parameter you have control on this \"feature\".",
      "votes": 1,
      "replies": [
        {
          "id": 759067,
          "postDate": "2020-02-28T14:26:22.060Z",
          "content": "<p>I think that wd greater than 1e-4 is dangerous. Thanks for the link, it was informative)</p>",
          "rawMarkdown": "I think that wd greater than 1e-4 is dangerous. Thanks for the link, it was informative)"
        },
        {
          "id": 839349,
          "postDate": "2020-05-09T10:36:25.620Z",
          "content": "<p>Why &gt; 1e-4 is harmful? Thanks for clarification. How to set this value? </p>",
          "rawMarkdown": "Why &gt; 1e-4 is harmful? Thanks for clarification. How to set this value? ",
          "isDeleted": true
        },
        {
          "id": 840404,
          "postDate": "2020-05-10T01:48:29.253Z",
          "content": "<p>Solved. I found one kernel to explain the weight decay and regularization\n<a href=\"https://www.kaggle.com/sid321axn/regularization-techniques-in-deep-learning\">https://www.kaggle.com/sid321axn/regularization-techniques-in-deep-learning</a></p>",
          "rawMarkdown": "Solved. I found one kernel to explain the weight decay and regularization\nhttps://www.kaggle.com/sid321axn/regularization-techniques-in-deep-learning\n",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 757336,
      "postDate": "2020-02-26T16:46:55.457Z",
      "content": "<p>Yes, it is. \nIt's very useful to decrees the optimizer steps size when you are approaching the minimum.\nOtherwise balancing very strong you can easily miss the minimum </p>",
      "rawMarkdown": "Yes, it is. \nIt's very useful to decrees the optimizer steps size when you are approaching the minimum.\nOtherwise balancing very strong you can easily miss the minimum ",
      "votes": 1,
      "replies": [
        {
          "id": 757396,
          "postDate": "2020-02-26T18:11:48.153Z",
          "content": "<p>Thanks. I`ll try it then)</p>",
          "rawMarkdown": "Thanks. I`ll try it then)"
        }
      ]
    },
    {
      "id": 757088,
      "postDate": "2020-02-26T12:23:53.343Z",
      "content": "<p>Is it helpful to use WD in Adam optimizer? I heard that it isn`t work correctly. I am using pytorch btw.</p>",
      "rawMarkdown": "Is it helpful to use WD in Adam optimizer? I heard that it isn`t work correctly. I am using pytorch btw.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 759075,
      "author_name": "ChaseWind",
      "author_url": "",
      "post_date": "2020-02-28T14:31:04.310000",
      "content": "<p>You should definitely check this paper: Decoupled Weight Decay Regularization ( <a href=\"https://arxiv.org/abs/1711.05101\">https://arxiv.org/abs/1711.05101</a> ).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 758926,
      "author_name": "Lasha Kharshiladze",
      "author_url": "",
      "post_date": "2020-02-28T10:56:09.660000",
      "content": "<p><a href=\"https://youtu.be/_JB0AO7QxSA?list=PLf7L7Kg8_FNxHATtLwDceyh72QQL9pvpQ&amp;t=1910\">here</a> Justin Johnson says that: there is theoritical work that most shallow minimums  that SGD-momentum misses are actually bad minimums - minimums that does not generalize well on test set, that we call local minimum, so because Adam optimizer is based on SGD momentum it also have that \"feature\", that it can jump over bad minimum,. So you can use regularization and don't care about landing in local minimum. Also for other side minimums that  SGD-momentum skips is not always bad minimums and it sometimes can be  best to land there so by regularization parameter you have control on this \"feature\".</p>",
      "votes": 1,
      "replies": [
        {
          "id": 759067,
          "author_name": "Kupchanski",
          "author_url": "",
          "post_date": "2020-02-28T14:26:22.060000",
          "content": "<p>I think that wd greater than 1e-4 is dangerous. Thanks for the link, it was informative)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 839349,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-09T10:36:25.620000",
          "content": "<p>Why &gt; 1e-4 is harmful? Thanks for clarification. How to set this value? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840404,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-10T01:48:29.253000",
          "content": "<p>Solved. I found one kernel to explain the weight decay and regularization\n<a href=\"https://www.kaggle.com/sid321axn/regularization-techniques-in-deep-learning\">https://www.kaggle.com/sid321axn/regularization-techniques-in-deep-learning</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757336,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-02-26T16:46:55.457000",
      "content": "<p>Yes, it is. \nIt's very useful to decrees the optimizer steps size when you are approaching the minimum.\nOtherwise balancing very strong you can easily miss the minimum </p>",
      "votes": 1,
      "replies": [
        {
          "id": 757396,
          "author_name": "Kupchanski",
          "author_url": "",
          "post_date": "2020-02-26T18:11:48.153000",
          "content": "<p>Thanks. I`ll try it then)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "759075": "You should definitely check this paper: Decoupled Weight Decay Regularization ( https://arxiv.org/abs/1711.05101 ).",
    "758926": "[here](https://youtu.be/_JB0AO7QxSA?list=PLf7L7Kg8_FNxHATtLwDceyh72QQL9pvpQ&amp;t=1910) Justin Johnson says that: there is theoritical work that most shallow minimums  that SGD-momentum misses are actually bad minimums - minimums that does not generalize well on test set, that we call local minimum, so because Adam optimizer is based on SGD momentum it also have that \"feature\", that it can jump over bad minimum,. So you can use regularization and don't care about landing in local minimum. Also for other side minimums that  SGD-momentum skips is not always bad minimums and it sometimes can be  best to land there so by regularization parameter you have control on this \"feature\".",
    "757336": "Yes, it is. \nIt's very useful to decrees the optimizer steps size when you are approaching the minimum.\nOtherwise balancing very strong you can easily miss the minimum ",
    "757088": "Is it helpful to use WD in Adam optimizer? I heard that it isn`t work correctly. I am using pytorch btw."
  }
}