{
  "id": 207560,
  "title": "Imbalanced Dataset - What should be do?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207560",
  "author_name": "",
  "post_date": "2020-12-30T08:24:51.567032700Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As we all know that the data set provided is highly imbalanced. </p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/50612932/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..5mD6tq9_fnO_e1sc0Uh6mA.LoxXeiJETMJmO_f-AMkae74a3JuWSBKSDrKsz99Q8JKH0AX8yhifuJWrolUiHeG0rFjH53c8ImenzdtjwMffgYmOLQaCL9zpgTzqQGG6XRewwrFtJNdIjDAEdG-NMJMP_2Ry9CU5sa86_o90YyNpjDqRzNqd0bkwpo_fwF7AJTd3vTrws16QWfdgjxVND-zKO3-BQzcpan1oM-F-5NhLv98zmOWjzLk3HECUSDDA8LZc9WFWGumweKKraqoguaz0G3nvUlbzbekGC3ao1WgeTfbvWa5bQ5HWsJj_4CIYyn6EvF3a_gh-vELeh-rHx21Bb8Shr41eEfW3gOQcjVRYuXb4SOzqzTKDJKPu7bQMIcdTa1tj9rdYoEP3Qq4MVEfJh_Ww7MxuIG_uklq_D8knONmYyUWTyQ85gFNShrTyRGQx0gp0vK9XeKheyzL00knX4qTxxdiZJ_srFlp9qvTY9UxfHG4A7WVUvrgb-F0H39DIJBMQaIwgZ1oRyRZKIJDNW7yDWMe-_AoF-dMq0I6OG7rC0Cfh1XbbsnPEcxtOj5o-VGa_YYTEqYeOBwyjiseCY3kvgASdWv-vMyw6rpfMvMCyW2gq2jJo1T9eedYZlnb-xs98-FMnnFmFvb0jGnw4TL-fhBVGSsRQoXLQCLbziddVwBFH7yx9kVcnyuUUK-M.CmWjeck0vZwfWN9L5PiNKg/__results___files/__results___11_0.png\" alt=\"Distribution Of Classes\"></p>\n<p>Shouldn't we do some oversampling or putting some class weight on the model?</p>",
  "messages": [
    {
      "id": "1132186",
      "postDate": "12/30/2020 08:24:51",
      "content": "<p>As we all know that the data set provided is highly imbalanced. </p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/50612932/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..5mD6tq9_fnO_e1sc0Uh6mA.LoxXeiJETMJmO_f-AMkae74a3JuWSBKSDrKsz99Q8JKH0AX8yhifuJWrolUiHeG0rFjH53c8ImenzdtjwMffgYmOLQaCL9zpgTzqQGG6XRewwrFtJNdIjDAEdG-NMJMP_2Ry9CU5sa86_o90YyNpjDqRzNqd0bkwpo_fwF7AJTd3vTrws16QWfdgjxVND-zKO3-BQzcpan1oM-F-5NhLv98zmOWjzLk3HECUSDDA8LZc9WFWGumweKKraqoguaz0G3nvUlbzbekGC3ao1WgeTfbvWa5bQ5HWsJj_4CIYyn6EvF3a_gh-vELeh-rHx21Bb8Shr41eEfW3gOQcjVRYuXb4SOzqzTKDJKPu7bQMIcdTa1tj9rdYoEP3Qq4MVEfJh_Ww7MxuIG_uklq_D8knONmYyUWTyQ85gFNShrTyRGQx0gp0vK9XeKheyzL00knX4qTxxdiZJ_srFlp9qvTY9UxfHG4A7WVUvrgb-F0H39DIJBMQaIwgZ1oRyRZKIJDNW7yDWMe-_AoF-dMq0I6OG7rC0Cfh1XbbsnPEcxtOj5o-VGa_YYTEqYeOBwyjiseCY3kvgASdWv-vMyw6rpfMvMCyW2gq2jJo1T9eedYZlnb-xs98-FMnnFmFvb0jGnw4TL-fhBVGSsRQoXLQCLbziddVwBFH7yx9kVcnyuUUK-M.CmWjeck0vZwfWN9L5PiNKg/__results___files/__results___11_0.png\" alt=\"Distribution Of Classes\"></p>\n<p>Shouldn't we do some oversampling or putting some class weight on the model?</p>",
      "rawMarkdown": "As we all know that the data set provided is highly imbalanced. \n\n![Distribution Of Classes](https://www.kaggleusercontent.com/kf/50612932/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..5mD6tq9_fnO_e1sc0Uh6mA.LoxXeiJETMJmO_f-AMkae74a3JuWSBKSDrKsz99Q8JKH0AX8yhifuJWrolUiHeG0rFjH53c8ImenzdtjwMffgYmOLQaCL9zpgTzqQGG6XRewwrFtJNdIjDAEdG-NMJMP_2Ry9CU5sa86_o90YyNpjDqRzNqd0bkwpo_fwF7AJTd3vTrws16QWfdgjxVND-zKO3-BQzcpan1oM-F-5NhLv98zmOWjzLk3HECUSDDA8LZc9WFWGumweKKraqoguaz0G3nvUlbzbekGC3ao1WgeTfbvWa5bQ5HWsJj_4CIYyn6EvF3a_gh-vELeh-rHx21Bb8Shr41eEfW3gOQcjVRYuXb4SOzqzTKDJKPu7bQMIcdTa1tj9rdYoEP3Qq4MVEfJh_Ww7MxuIG_uklq_D8knONmYyUWTyQ85gFNShrTyRGQx0gp0vK9XeKheyzL00knX4qTxxdiZJ_srFlp9qvTY9UxfHG4A7WVUvrgb-F0H39DIJBMQaIwgZ1oRyRZKIJDNW7yDWMe-_AoF-dMq0I6OG7rC0Cfh1XbbsnPEcxtOj5o-VGa_YYTEqYeOBwyjiseCY3kvgASdWv-vMyw6rpfMvMCyW2gq2jJo1T9eedYZlnb-xs98-FMnnFmFvb0jGnw4TL-fhBVGSsRQoXLQCLbziddVwBFH7yx9kVcnyuUUK-M.CmWjeck0vZwfWN9L5PiNKg/__results___files/__results___11_0.png)\n\nShouldn't we do some oversampling or putting some class weight on the model?",
      "votes": null
    },
    {
      "id": "1132191",
      "postDate": "12/30/2020 08:27:11",
      "content": "<p>I tried putting some class weights on my model. But it did not perform well and the accuracy dropped. Any comment why is it so? or what should be do for imbalanced datasets?<br>\n<a href=\"https://www.kaggle.com/shubham219/experiment-with-models-using-keras-with-updates\" target=\"_blank\">My Notebook</a></p>",
      "rawMarkdown": "I tried putting some class weights on my model. But it did not perform well and the accuracy dropped. Any comment why is it so? or what should be do for imbalanced datasets?\n[My Notebook](https://www.kaggle.com/shubham219/experiment-with-models-using-keras-with-updates)",
      "votes": null
    },
    {
      "id": "1132881",
      "postDate": "12/30/2020 18:58:30",
      "content": "<p>There is a good paper I would like to recommend, it covers exactly this problem.</p>\n<p>Quote:<br>\n<code>In the  context  of  deep  feature  representation  learning  using CNNs, re-sampling may either introduce large amounts of duplicated  samples,  which  slows  down  the  training  and makes  the  model  susceptible  to  overfitting  when  over-sampling, or discard valuable examples that are important for feature learning when under-sampling.</code></p>\n<p>Regarding this work, re-sampling is a worse method than re-weighting the loss function. They also claim that undersampling is generally considered to be better than oversampling.</p>\n<p>Paper-Link: <a href=\"https://arxiv.org/pdf/1901.05555.pdf\" target=\"_blank\">https://arxiv.org/pdf/1901.05555.pdf</a><br>\nImplementation of their loss function: <a href=\"https://github.com/vandit15/Class-balanced-loss-pytorch\" target=\"_blank\">https://github.com/vandit15/Class-balanced-loss-pytorch</a></p>",
      "rawMarkdown": "There is a good paper I would like to recommend, it covers exactly this problem.\n\nQuote:\n`In the  context  of  deep  feature  representation  learning  using CNNs, re-sampling may either introduce large amounts of duplicated  samples,  which  slows  down  the  training  and makes  the  model  susceptible  to  overfitting  when  over-sampling, or discard valuable examples that are important for feature learning when under-sampling. `\n\nRegarding this work, re-sampling is a worse method than re-weighting the loss function. They also claim that undersampling is generally considered to be better than oversampling.\n\nPaper-Link: https://arxiv.org/pdf/1901.05555.pdf\nImplementation of their loss function: https://github.com/vandit15/Class-balanced-loss-pytorch",
      "votes": null
    },
    {
      "id": "1133331",
      "postDate": "12/31/2020 05:50:07",
      "content": "<p>Thank You! for sharing this I will try and get back on this post again</p>",
      "rawMarkdown": "Thank You! for sharing this I will try and get back on this post again",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1132191,
      "author_name": "shubham219",
      "author_url": "",
      "post_date": "12/30/2020 08:27:11",
      "content": "<p>I tried putting some class weights on my model. But it did not perform well and the accuracy dropped. Any comment why is it so? or what should be do for imbalanced datasets?<br>\n<a href=\"https://www.kaggle.com/shubham219/experiment-with-models-using-keras-with-updates\" target=\"_blank\">My Notebook</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1132881,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "12/30/2020 18:58:30",
      "content": "<p>There is a good paper I would like to recommend, it covers exactly this problem.</p>\n<p>Quote:<br>\n<code>In the  context  of  deep  feature  representation  learning  using CNNs, re-sampling may either introduce large amounts of duplicated  samples,  which  slows  down  the  training  and makes  the  model  susceptible  to  overfitting  when  over-sampling, or discard valuable examples that are important for feature learning when under-sampling.</code></p>\n<p>Regarding this work, re-sampling is a worse method than re-weighting the loss function. They also claim that undersampling is generally considered to be better than oversampling.</p>\n<p>Paper-Link: <a href=\"https://arxiv.org/pdf/1901.05555.pdf\" target=\"_blank\">https://arxiv.org/pdf/1901.05555.pdf</a><br>\nImplementation of their loss function: <a href=\"https://github.com/vandit15/Class-balanced-loss-pytorch\" target=\"_blank\">https://github.com/vandit15/Class-balanced-loss-pytorch</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1133331,
          "author_name": "shubham219",
          "author_url": "",
          "post_date": "12/31/2020 05:50:07",
          "content": "<p>Thank You! for sharing this I will try and get back on this post again</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1132186": "As we all know that the data set provided is highly imbalanced. \n\n![Distribution Of Classes](https://www.kaggleusercontent.com/kf/50612932/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..5mD6tq9_fnO_e1sc0Uh6mA.LoxXeiJETMJmO_f-AMkae74a3JuWSBKSDrKsz99Q8JKH0AX8yhifuJWrolUiHeG0rFjH53c8ImenzdtjwMffgYmOLQaCL9zpgTzqQGG6XRewwrFtJNdIjDAEdG-NMJMP_2Ry9CU5sa86_o90YyNpjDqRzNqd0bkwpo_fwF7AJTd3vTrws16QWfdgjxVND-zKO3-BQzcpan1oM-F-5NhLv98zmOWjzLk3HECUSDDA8LZc9WFWGumweKKraqoguaz0G3nvUlbzbekGC3ao1WgeTfbvWa5bQ5HWsJj_4CIYyn6EvF3a_gh-vELeh-rHx21Bb8Shr41eEfW3gOQcjVRYuXb4SOzqzTKDJKPu7bQMIcdTa1tj9rdYoEP3Qq4MVEfJh_Ww7MxuIG_uklq_D8knONmYyUWTyQ85gFNShrTyRGQx0gp0vK9XeKheyzL00knX4qTxxdiZJ_srFlp9qvTY9UxfHG4A7WVUvrgb-F0H39DIJBMQaIwgZ1oRyRZKIJDNW7yDWMe-_AoF-dMq0I6OG7rC0Cfh1XbbsnPEcxtOj5o-VGa_YYTEqYeOBwyjiseCY3kvgASdWv-vMyw6rpfMvMCyW2gq2jJo1T9eedYZlnb-xs98-FMnnFmFvb0jGnw4TL-fhBVGSsRQoXLQCLbziddVwBFH7yx9kVcnyuUUK-M.CmWjeck0vZwfWN9L5PiNKg/__results___files/__results___11_0.png)\n\nShouldn't we do some oversampling or putting some class weight on the model?",
    "1132191": "I tried putting some class weights on my model. But it did not perform well and the accuracy dropped. Any comment why is it so? or what should be do for imbalanced datasets?\n[My Notebook](https://www.kaggle.com/shubham219/experiment-with-models-using-keras-with-updates)",
    "1132881": "There is a good paper I would like to recommend, it covers exactly this problem.\n\nQuote:\n`In the  context  of  deep  feature  representation  learning  using CNNs, re-sampling may either introduce large amounts of duplicated  samples,  which  slows  down  the  training  and makes  the  model  susceptible  to  overfitting  when  over-sampling, or discard valuable examples that are important for feature learning when under-sampling. `\n\nRegarding this work, re-sampling is a worse method than re-weighting the loss function. They also claim that undersampling is generally considered to be better than oversampling.\n\nPaper-Link: https://arxiv.org/pdf/1901.05555.pdf\nImplementation of their loss function: https://github.com/vandit15/Class-balanced-loss-pytorch",
    "1133331": "Thank You! for sharing this I will try and get back on this post again"
  },
  "source": "meta"
}