{
  "id": 130097,
  "title": "How do you cope with overfitting?",
  "url": "/competitions/bengaliai-cv19/discussion/130097",
  "author_name": "Volodymyr",
  "post_date": "2020-02-12T06:42:15.375000",
  "votes": 11,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have tried se_resnext50_32x4d and se_resnext101_32x4d and for them I end up training by early stopping on 16 and 30 epochs respectively. While most of the participants reports training for 100 epochs\nI use early stopping by CrossEntropy, Adam(1e-4 - base model and 5e-4 classifier), ReduceOnPlateau. Also I use AugMix (severity=2) and Dropout for clf (p=0.2)\nAlso for se_resnext101_32x4d I noticed that final train CE is bigger 10 times than final validation CE!\nDoes anyone has came across with the same problem? Or maybe is there any other tricks how to cope with overfitting? </p>",
  "messages": [
    {
      "id": 743634,
      "postDate": "2020-02-12T06:42:15.377Z",
      "content": "<p>I have tried se_resnext50_32x4d and se_resnext101_32x4d and for them I end up training by early stopping on 16 and 30 epochs respectively. While most of the participants reports training for 100 epochs\nI use early stopping by CrossEntropy, Adam(1e-4 - base model and 5e-4 classifier), ReduceOnPlateau. Also I use AugMix (severity=2) and Dropout for clf (p=0.2)\nAlso for se_resnext101_32x4d I noticed that final train CE is bigger 10 times than final validation CE!\nDoes anyone has came across with the same problem? Or maybe is there any other tricks how to cope with overfitting? </p>",
      "rawMarkdown": "I have tried se_resnext50_32x4d and se_resnext101_32x4d and for them I end up training by early stopping on 16 and 30 epochs respectively. While most of the participants reports training for 100 epochs\nI use early stopping by CrossEntropy, Adam(1e-4 - base model and 5e-4 classifier), ReduceOnPlateau. Also I use AugMix (severity=2) and Dropout for clf (p=0.2)\nAlso for se_resnext101_32x4d I noticed that final train CE is bigger 10 times than final validation CE!\nDoes anyone has came across with the same problem? Or maybe is there any other tricks how to cope with overfitting? ",
      "votes": 11
    },
    {
      "id": 747189,
      "postDate": "2020-02-16T05:02:10.073Z",
      "content": "<p>maybe overfitting will become a thing of the past</p>\n\n<p><img src=\"https://pbs.twimg.com/media/EQtm-o_VAAQwtXj?format=jpg&amp;name=4096x4096\" alt=\"\"></p>\n\n<p><a href=\"https://ipam.wistia.com/medias/w1ii9wdqnd\">https://ipam.wistia.com/medias/w1ii9wdqnd</a></p>",
      "rawMarkdown": "maybe overfitting will become a thing of the past\n\n![](https://pbs.twimg.com/media/EQtm-o_VAAQwtXj?format=jpg&amp;name=4096x4096)\n\nhttps://ipam.wistia.com/medias/w1ii9wdqnd",
      "votes": 1,
      "replies": [
        {
          "id": 764337,
          "postDate": "2020-03-05T11:08:55.750Z",
          "content": "<p>maybe through better regularization :)\n</p>\n\n<p><strong>Arxiv</strong>: <a href=\"https://arxiv.org/abs/2003.01897\">https://arxiv.org/abs/2003.01897</a>\n<strong>Twitter thread</strong>: <a href=\"https://twitter.com/PreetumNakkiran/status/1235376866715820032\">https://twitter.com/PreetumNakkiran/status/1235376866715820032</a></p>",
          "rawMarkdown": "maybe through better regularization :)\n<img src=\"https://pbs.twimg.com/media/ESTwNyOU8AA4Mio?format=jpg\" width=\"650\">\n\n**Arxiv**: https://arxiv.org/abs/2003.01897\n**Twitter thread**: https://twitter.com/PreetumNakkiran/status/1235376866715820032",
          "votes": 1
        }
      ]
    },
    {
      "id": 744309,
      "postDate": "2020-02-12T18:10:50.740Z",
      "content": "<p>For me AugMix didn't give any boost on the top of my current best augmentations (Cutout + some usual augs == Cutout + some usual augs + AugMix).\nSolo AugMix barely achieved CV 0.970 ( LB 0.960) and converges already in 20-30 epochs.</p>",
      "rawMarkdown": "For me AugMix didn't give any boost on the top of my current best augmentations (Cutout + some usual augs == Cutout + some usual augs + AugMix).\nSolo AugMix barely achieved CV 0.970 ( LB 0.960) and converges already in 20-30 epochs.",
      "votes": 1,
      "replies": [
        {
          "id": 744425,
          "postDate": "2020-02-12T20:31:33.940Z",
          "content": "<p>Hm, sounds that augmix is not strong enough for this dataset</p>",
          "rawMarkdown": "Hm, sounds that augmix is not strong enough for this dataset",
          "votes": 1
        },
        {
          "id": 744428,
          "postDate": "2020-02-12T20:34:49.743Z",
          "content": "<p>By the way you also have a 1% gap between CV and LB score?</p>",
          "rawMarkdown": "By the way you also have a 1% gap between CV and LB score?",
          "votes": 1
        },
        {
          "id": 744892,
          "postDate": "2020-02-13T09:24:21.147Z",
          "content": "<p>Yes, this is right for almost all my experiments: 1-1.1% gap between CV and LB. But I always check only one fold, no assembles.</p>",
          "rawMarkdown": "Yes, this is right for almost all my experiments: 1-1.1% gap between CV and LB. But I always check only one fold, no assembles.",
          "votes": 1
        }
      ]
    },
    {
      "id": 744136,
      "postDate": "2020-02-12T15:40:31.203Z",
      "content": "<p>Well, if you are going with seresnext10132x4d you have to be sure that:\na) you don't add up too many layers/neurons after the classifier (being dense layers, the weights are adding up fast and you model will be to big for our problem)\nb) you should augment pretty hardcore (either cutmix/mixup or affine transforms + cutout) \nc) add regularization</p>\n\n<p>By the way, what resolution are you using for the images ?</p>",
      "rawMarkdown": "Well, if you are going with seresnext10132x4d you have to be sure that:\na) you don't add up too many layers/neurons after the classifier (being dense layers, the weights are adding up fast and you model will be to big for our problem)\nb) you should augment pretty hardcore (either cutmix/mixup or affine transforms + cutout) \nc) add regularization\n\nBy the way, what resolution are you using for the images ?",
      "votes": 1,
      "replies": [
        {
          "id": 744427,
          "postDate": "2020-02-12T20:34:00.803Z",
          "content": "<p>a) I have only one Linear after seresnext10132x4d\nb) yep. It seems that augmix was not enough and adding mixups and cutmixes is a good wayout\nc) I will try\nI am using original resolution\nThanks for your advice!</p>",
          "rawMarkdown": "a) I have only one Linear after seresnext10132x4d\nb) yep. It seems that augmix was not enough and adding mixups and cutmixes is a good wayout\nc) I will try\nI am using original resolution\nThanks for your advice!"
        }
      ]
    },
    {
      "id": 743693,
      "postDate": "2020-02-12T07:47:11.613Z",
      "content": "<p>Cutmix, Mixup, intensive regularization techniques like Shake drop, and heavier augmentations will help reduce overfitting </p>",
      "rawMarkdown": "Cutmix, Mixup, intensive regularization techniques like Shake drop, and heavier augmentations will help reduce overfitting "
    },
    {
      "id": 745834,
      "postDate": "2020-02-14T09:04:18.183Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 747189,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-16T05:02:10.073000",
      "content": "<p>maybe overfitting will become a thing of the past</p>\n\n<p><img src=\"https://pbs.twimg.com/media/EQtm-o_VAAQwtXj?format=jpg&amp;name=4096x4096\" alt=\"\"></p>\n\n<p><a href=\"https://ipam.wistia.com/medias/w1ii9wdqnd\">https://ipam.wistia.com/medias/w1ii9wdqnd</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 764337,
          "author_name": "Onur Tunali",
          "author_url": "",
          "post_date": "2020-03-05T11:08:55.750000",
          "content": "<p>maybe through better regularization :)\n</p>\n\n<p><strong>Arxiv</strong>: <a href=\"https://arxiv.org/abs/2003.01897\">https://arxiv.org/abs/2003.01897</a>\n<strong>Twitter thread</strong>: <a href=\"https://twitter.com/PreetumNakkiran/status/1235376866715820032\">https://twitter.com/PreetumNakkiran/status/1235376866715820032</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 744309,
      "author_name": "Andrei Dukhounik",
      "author_url": "",
      "post_date": "2020-02-12T18:10:50.740000",
      "content": "<p>For me AugMix didn't give any boost on the top of my current best augmentations (Cutout + some usual augs == Cutout + some usual augs + AugMix).\nSolo AugMix barely achieved CV 0.970 ( LB 0.960) and converges already in 20-30 epochs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 744425,
          "author_name": "Volodymyr",
          "author_url": "",
          "post_date": "2020-02-12T20:31:33.940000",
          "content": "<p>Hm, sounds that augmix is not strong enough for this dataset</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744428,
          "author_name": "Volodymyr",
          "author_url": "",
          "post_date": "2020-02-12T20:34:49.743000",
          "content": "<p>By the way you also have a 1% gap between CV and LB score?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744892,
          "author_name": "Andrei Dukhounik",
          "author_url": "",
          "post_date": "2020-02-13T09:24:21.147000",
          "content": "<p>Yes, this is right for almost all my experiments: 1-1.1% gap between CV and LB. But I always check only one fold, no assembles.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 744136,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-02-12T15:40:31.203000",
      "content": "<p>Well, if you are going with seresnext10132x4d you have to be sure that:\na) you don't add up too many layers/neurons after the classifier (being dense layers, the weights are adding up fast and you model will be to big for our problem)\nb) you should augment pretty hardcore (either cutmix/mixup or affine transforms + cutout) \nc) add regularization</p>\n\n<p>By the way, what resolution are you using for the images ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 744427,
          "author_name": "Volodymyr",
          "author_url": "",
          "post_date": "2020-02-12T20:34:00.803000",
          "content": "<p>a) I have only one Linear after seresnext10132x4d\nb) yep. It seems that augmix was not enough and adding mixups and cutmixes is a good wayout\nc) I will try\nI am using original resolution\nThanks for your advice!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 743693,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2020-02-12T07:47:11.613000",
      "content": "<p>Cutmix, Mixup, intensive regularization techniques like Shake drop, and heavier augmentations will help reduce overfitting </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 745834,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-14T09:04:18.183000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "743634": "I have tried se_resnext50_32x4d and se_resnext101_32x4d and for them I end up training by early stopping on 16 and 30 epochs respectively. While most of the participants reports training for 100 epochs\nI use early stopping by CrossEntropy, Adam(1e-4 - base model and 5e-4 classifier), ReduceOnPlateau. Also I use AugMix (severity=2) and Dropout for clf (p=0.2)\nAlso for se_resnext101_32x4d I noticed that final train CE is bigger 10 times than final validation CE!\nDoes anyone has came across with the same problem? Or maybe is there any other tricks how to cope with overfitting? ",
    "747189": "maybe overfitting will become a thing of the past\n\n![](https://pbs.twimg.com/media/EQtm-o_VAAQwtXj?format=jpg&amp;name=4096x4096)\n\nhttps://ipam.wistia.com/medias/w1ii9wdqnd",
    "744309": "For me AugMix didn't give any boost on the top of my current best augmentations (Cutout + some usual augs == Cutout + some usual augs + AugMix).\nSolo AugMix barely achieved CV 0.970 ( LB 0.960) and converges already in 20-30 epochs.",
    "744136": "Well, if you are going with seresnext10132x4d you have to be sure that:\na) you don't add up too many layers/neurons after the classifier (being dense layers, the weights are adding up fast and you model will be to big for our problem)\nb) you should augment pretty hardcore (either cutmix/mixup or affine transforms + cutout) \nc) add regularization\n\nBy the way, what resolution are you using for the images ?",
    "743693": "Cutmix, Mixup, intensive regularization techniques like Shake drop, and heavier augmentations will help reduce overfitting ",
    "745834": ""
  }
}