{
  "id": 242455,
  "title": "How did you handle class imbalance in the dataset?",
  "url": "/competitions/iwildcam2021-fgvc8/discussion/242455",
  "author_name": "",
  "post_date": "2021-05-29T04:35:32.441875600Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It would be so amazing to know about various dataset imbalance techniques you guys might have used. Which techniques did work for you? And what were your experiments?</p>\n<p>My approach - <br>\nI created a small training and validation subset (approx 10% of the actual dataset), and used a small model \"tf_efficientnet_b2_ns\" to try out following techniques - </p>\n<ol>\n<li><p>Use weighted cross-entropy loss<br>\nWe can assign weights to the cross-entropy loss such that it will penalize more to the smaller classes and the less to larger classes. <a href=\"https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596\" target=\"_blank\">Here</a> is how I did it in PyTorch.</p></li>\n<li><p>Use focal loss<br>\nOriginally proposed for object detection, but we can also use this for any other use cases. You can read more about it <a href=\"https://amaarora.github.io/2020/06/29/FocalLoss.html\" target=\"_blank\">here</a>. I used this <a href=\"https://github.com/AdeelH/pytorch-multi-class-focal-loss\" target=\"_blank\">one</a>.</p></li>\n<li><p>Over Sampling and Under Sampling<br>\nI used imblearn and randomly upsampled classes smaller than 25 to 25, and randomly down sampled classes larger than 7000 to 7000. As a result, any class having less than 25 instances got upsampled to 25 and any class having more than 7000 instances got downsampled to 7000. I kept in mind that the competition test set will also be imbalanced, so I didn't want to correct imbalance totally. </p></li>\n<li><p>Create a separate model for small classes<br>\nWe had some classes that had very small number of instances, so I created a separate classifier for these small classes (called small_classifier for eg). I grouped together these small clases under a single class (called small_class for eg) so that my main classifier will classify small_class with all other big classes in the dataset. And if my main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. I expected this technique to give me accuracy boosts as now main classifier does not need to deal with very small classes, and instead small_classifier will be looking only at these small classes.</p></li>\n</ol>\n<p>At the end, combining technique 1 and 3 performed the best on my small validation set. Hence I have used the same for my final submission.</p>\n<p>Would love to hear how you guys tackled this issue. :) </p>",
  "messages": [
    {
      "id": "1327216",
      "postDate": "05/29/2021 04:35:32",
      "content": "<p>It would be so amazing to know about various dataset imbalance techniques you guys might have used. Which techniques did work for you? And what were your experiments?</p>\n<p>My approach - <br>\nI created a small training and validation subset (approx 10% of the actual dataset), and used a small model \"tf_efficientnet_b2_ns\" to try out following techniques - </p>\n<ol>\n<li><p>Use weighted cross-entropy loss<br>\nWe can assign weights to the cross-entropy loss such that it will penalize more to the smaller classes and the less to larger classes. <a href=\"https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596\" target=\"_blank\">Here</a> is how I did it in PyTorch.</p></li>\n<li><p>Use focal loss<br>\nOriginally proposed for object detection, but we can also use this for any other use cases. You can read more about it <a href=\"https://amaarora.github.io/2020/06/29/FocalLoss.html\" target=\"_blank\">here</a>. I used this <a href=\"https://github.com/AdeelH/pytorch-multi-class-focal-loss\" target=\"_blank\">one</a>.</p></li>\n<li><p>Over Sampling and Under Sampling<br>\nI used imblearn and randomly upsampled classes smaller than 25 to 25, and randomly down sampled classes larger than 7000 to 7000. As a result, any class having less than 25 instances got upsampled to 25 and any class having more than 7000 instances got downsampled to 7000. I kept in mind that the competition test set will also be imbalanced, so I didn't want to correct imbalance totally. </p></li>\n<li><p>Create a separate model for small classes<br>\nWe had some classes that had very small number of instances, so I created a separate classifier for these small classes (called small_classifier for eg). I grouped together these small clases under a single class (called small_class for eg) so that my main classifier will classify small_class with all other big classes in the dataset. And if my main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. I expected this technique to give me accuracy boosts as now main classifier does not need to deal with very small classes, and instead small_classifier will be looking only at these small classes.</p></li>\n</ol>\n<p>At the end, combining technique 1 and 3 performed the best on my small validation set. Hence I have used the same for my final submission.</p>\n<p>Would love to hear how you guys tackled this issue. :) </p>",
      "rawMarkdown": "It would be so amazing to know about various dataset imbalance techniques you guys might have used. Which techniques did work for you? And what were your experiments?\n\nMy approach - \nI created a small training and validation subset (approx 10% of the actual dataset), and used a small model \"tf_efficientnet_b2_ns\" to try out following techniques - \n\n1. Use weighted cross-entropy loss\nWe can assign weights to the cross-entropy loss such that it will penalize more to the smaller classes and the less to larger classes. [Here](https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596) is how I did it in PyTorch.\n\n2. Use focal loss\nOriginally proposed for object detection, but we can also use this for any other use cases. You can read more about it [here](https://amaarora.github.io/2020/06/29/FocalLoss.html). I used this [one](https://github.com/AdeelH/pytorch-multi-class-focal-loss).\n\n3. Over Sampling and Under Sampling\nI used imblearn and randomly upsampled classes smaller than 25 to 25, and randomly down sampled classes larger than 7000 to 7000. As a result, any class having less than 25 instances got upsampled to 25 and any class having more than 7000 instances got downsampled to 7000. I kept in mind that the competition test set will also be imbalanced, so I didn't want to correct imbalance totally. \n\n4. Create a separate model for small classes\nWe had some classes that had very small number of instances, so I created a separate classifier for these small classes (called small_classifier for eg). I grouped together these small clases under a single class (called small_class for eg) so that my main classifier will classify small_class with all other big classes in the dataset. And if my main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. I expected this technique to give me accuracy boosts as now main classifier does not need to deal with very small classes, and instead small_classifier will be looking only at these small classes.\n\nAt the end, combining technique 1 and 3 performed the best on my small validation set. Hence I have used the same for my final submission.\n\nWould love to hear how you guys tackled this issue. :)",
      "votes": null
    },
    {
      "id": "1344529",
      "postDate": "06/11/2021 01:51:32",
      "content": "<p>For our solution, we used <a href=\"https://arxiv.org/abs/2006.10408\" target=\"_blank\">Balanced Group Softmax</a>.</p>",
      "rawMarkdown": "For our solution, we used [Balanced Group Softmax](https://arxiv.org/abs/2006.10408).",
      "votes": null
    },
    {
      "id": "1344727",
      "postDate": "06/11/2021 05:18:24",
      "content": "<p>Wow!! amazing paper, thank you for sharing.</p>\n<p>And I am very eager to know about your solution, it would be so amazing.</p>",
      "rawMarkdown": "Wow!! amazing paper, thank you for sharing.\n\nAnd I am very eager to know about your solution, it would be so amazing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1344529,
      "author_name": "fagner",
      "author_url": "",
      "post_date": "06/11/2021 01:51:32",
      "content": "<p>For our solution, we used <a href=\"https://arxiv.org/abs/2006.10408\" target=\"_blank\">Balanced Group Softmax</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1344727,
          "author_name": "devashishprasad",
          "author_url": "",
          "post_date": "06/11/2021 05:18:24",
          "content": "<p>Wow!! amazing paper, thank you for sharing.</p>\n<p>And I am very eager to know about your solution, it would be so amazing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1327216": "It would be so amazing to know about various dataset imbalance techniques you guys might have used. Which techniques did work for you? And what were your experiments?\n\nMy approach - \nI created a small training and validation subset (approx 10% of the actual dataset), and used a small model \"tf_efficientnet_b2_ns\" to try out following techniques - \n\n1. Use weighted cross-entropy loss\nWe can assign weights to the cross-entropy loss such that it will penalize more to the smaller classes and the less to larger classes. [Here](https://discuss.pytorch.org/t/dealing-with-imbalanced-datasets-in-pytorch/22596) is how I did it in PyTorch.\n\n2. Use focal loss\nOriginally proposed for object detection, but we can also use this for any other use cases. You can read more about it [here](https://amaarora.github.io/2020/06/29/FocalLoss.html). I used this [one](https://github.com/AdeelH/pytorch-multi-class-focal-loss).\n\n3. Over Sampling and Under Sampling\nI used imblearn and randomly upsampled classes smaller than 25 to 25, and randomly down sampled classes larger than 7000 to 7000. As a result, any class having less than 25 instances got upsampled to 25 and any class having more than 7000 instances got downsampled to 7000. I kept in mind that the competition test set will also be imbalanced, so I didn't want to correct imbalance totally. \n\n4. Create a separate model for small classes\nWe had some classes that had very small number of instances, so I created a separate classifier for these small classes (called small_classifier for eg). I grouped together these small clases under a single class (called small_class for eg) so that my main classifier will classify small_class with all other big classes in the dataset. And if my main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. I expected this technique to give me accuracy boosts as now main classifier does not need to deal with very small classes, and instead small_classifier will be looking only at these small classes.\n\nAt the end, combining technique 1 and 3 performed the best on my small validation set. Hence I have used the same for my final submission.\n\nWould love to hear how you guys tackled this issue. :)",
    "1344529": "For our solution, we used [Balanced Group Softmax](https://arxiv.org/abs/2006.10408).",
    "1344727": "Wow!! amazing paper, thank you for sharing.\n\nAnd I am very eager to know about your solution, it would be so amazing."
  },
  "source": "meta"
}