{
  "id": 203352,
  "title": "Beginner Question about Weight Decay Regularization",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/203352",
  "author_name": "",
  "post_date": "2020-12-14T22:39:45.271288700Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm new to Kaggle (and any practical form of Data Science) but I've taken a few ML courses. In the classroom we learned a lot about weight decay but I haven't seen it discussed in this competition or explicitly implemented in public notebooks. Is it not necessary given the other, perhaps more advanced, forms of regularization such as Cutmix? Or is it being implemented under the hood? I've noticed in Keras that we need to add weight decay layer by layer so I'm thinking it was maybe done already in the original ImageNet model?</p>",
  "messages": [
    {
      "id": "1112799",
      "postDate": "12/14/2020 22:39:45",
      "content": "<p>I'm new to Kaggle (and any practical form of Data Science) but I've taken a few ML courses. In the classroom we learned a lot about weight decay but I haven't seen it discussed in this competition or explicitly implemented in public notebooks. Is it not necessary given the other, perhaps more advanced, forms of regularization such as Cutmix? Or is it being implemented under the hood? I've noticed in Keras that we need to add weight decay layer by layer so I'm thinking it was maybe done already in the original ImageNet model?</p>",
      "rawMarkdown": "I'm new to Kaggle (and any practical form of Data Science) but I've taken a few ML courses. In the classroom we learned a lot about weight decay but I haven't seen it discussed in this competition or explicitly implemented in public notebooks. Is it not necessary given the other, perhaps more advanced, forms of regularization such as Cutmix? Or is it being implemented under the hood? I've noticed in Keras that we need to add weight decay layer by layer so I'm thinking it was maybe done already in the original ImageNet model?",
      "votes": null
    },
    {
      "id": "1115223",
      "postDate": "12/16/2020 05:22:35",
      "content": "<p>It's used in optimizers as a parameter, no way for to Regularization be absent!</p>",
      "rawMarkdown": "It's used in optimizers as a parameter, no way for to Regularization be absent!",
      "votes": null
    },
    {
      "id": "1115226",
      "postDate": "12/16/2020 05:27:08",
      "content": "<p>judging from heatmap, we can find out that cutmix empowers models to identify different classes even if inside single image, by local features. This is very efficient for training.</p>",
      "rawMarkdown": "judging from heatmap, we can find out that cutmix empowers models to identify different classes even if inside single image, by local features. This is very efficient for training.",
      "votes": null
    },
    {
      "id": "1117446",
      "postDate": "12/18/2020 04:12:37",
      "content": "<p>Ah. I see that it is used in the tensorflow addon AdamW optimizer. But I haven't seen it in the standard optimizers. Are there any standard optimizers that implement it?</p>\n<p><a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/AdamW\" target=\"_blank\">https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/AdamW</a></p>",
      "rawMarkdown": "Ah. I see that it is used in the tensorflow addon AdamW optimizer. But I haven't seen it in the standard optimizers. Are there any standard optimizers that implement it?\n\nhttps://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/AdamW",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1115223,
      "author_name": "plugin1689",
      "author_url": "",
      "post_date": "12/16/2020 05:22:35",
      "content": "<p>It's used in optimizers as a parameter, no way for to Regularization be absent!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1117446,
          "author_name": "kagglethomas88",
          "author_url": "",
          "post_date": "12/18/2020 04:12:37",
          "content": "<p>Ah. I see that it is used in the tensorflow addon AdamW optimizer. But I haven't seen it in the standard optimizers. Are there any standard optimizers that implement it?</p>\n<p><a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/AdamW\" target=\"_blank\">https://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/AdamW</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115226,
      "author_name": "plugin1689",
      "author_url": "",
      "post_date": "12/16/2020 05:27:08",
      "content": "<p>judging from heatmap, we can find out that cutmix empowers models to identify different classes even if inside single image, by local features. This is very efficient for training.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1112799": "I'm new to Kaggle (and any practical form of Data Science) but I've taken a few ML courses. In the classroom we learned a lot about weight decay but I haven't seen it discussed in this competition or explicitly implemented in public notebooks. Is it not necessary given the other, perhaps more advanced, forms of regularization such as Cutmix? Or is it being implemented under the hood? I've noticed in Keras that we need to add weight decay layer by layer so I'm thinking it was maybe done already in the original ImageNet model?",
    "1115223": "It's used in optimizers as a parameter, no way for to Regularization be absent!",
    "1115226": "judging from heatmap, we can find out that cutmix empowers models to identify different classes even if inside single image, by local features. This is very efficient for training.",
    "1117446": "Ah. I see that it is used in the tensorflow addon AdamW optimizer. But I haven't seen it in the standard optimizers. Are there any standard optimizers that implement it?\n\nhttps://www.tensorflow.org/addons/api_docs/python/tfa/optimizers/AdamW"
  },
  "source": "meta"
}