{
  "id": 496556,
  "title": "Does handling class imbalance helps ?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/496556",
  "author_name": "Anuj Panthri",
  "post_date": "2024-04-21T15:12:48.137000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello everyone I'm a beginner, I was doing this challenge recently and I tried to make a simple approach and I found the following results :-</p>\n<ol>\n<li><strong>baseline</strong>(ResNet50_256_image_size_30_epochs): (public/private) (0.7605/0.7483)</li>\n<li>baseline+<strong>class_weights</strong>(class weights for class imbalance): doesn't helps (0.7375/0.7542)</li>\n<li>baseline+<strong>512_image_size</strong>: does helps (public/private) (0.7875/0.7938)</li>\n<li>baseline+<strong>label_smoothing</strong>(label smoothing handling training with noisy data): does helps too (public/private) (0.7718/0.7748)</li>\n<li>baseline+label_smoothing+<strong>data_augmentation+dropout+60_epochs</strong>: does helps too (0.8253/0.8383)</li>\n</ol>\n<p><strong>Note that in all these experiments I saw overfitting .</strong></p>\n<p>As you can see that using class_weights for handling class imbalance helped with private score on LB but it performed worse on the public score on LB.</p>\n<p>I read other's results and by what I understood was that they improved their results using center_cropping, TTA, using data augmentations, ensembling , but only a few people tried to focus on the class imbalance problem , is it not important in this case ? probably because of the class imbalance is also present in the test set and there are mislabeled examples ?</p>\n<p>Did it helped others to handle the class imbalance problem ? If yes then how much ?</p>",
  "messages": [
    {
      "id": 2766248,
      "postDate": "2024-04-21T15:12:48.137Z",
      "content": "<p>Hello everyone I'm a beginner, I was doing this challenge recently and I tried to make a simple approach and I found the following results :-</p>\n<ol>\n<li><strong>baseline</strong>(ResNet50_256_image_size_30_epochs): (public/private) (0.7605/0.7483)</li>\n<li>baseline+<strong>class_weights</strong>(class weights for class imbalance): doesn't helps (0.7375/0.7542)</li>\n<li>baseline+<strong>512_image_size</strong>: does helps (public/private) (0.7875/0.7938)</li>\n<li>baseline+<strong>label_smoothing</strong>(label smoothing handling training with noisy data): does helps too (public/private) (0.7718/0.7748)</li>\n<li>baseline+label_smoothing+<strong>data_augmentation+dropout+60_epochs</strong>: does helps too (0.8253/0.8383)</li>\n</ol>\n<p><strong>Note that in all these experiments I saw overfitting .</strong></p>\n<p>As you can see that using class_weights for handling class imbalance helped with private score on LB but it performed worse on the public score on LB.</p>\n<p>I read other's results and by what I understood was that they improved their results using center_cropping, TTA, using data augmentations, ensembling , but only a few people tried to focus on the class imbalance problem , is it not important in this case ? probably because of the class imbalance is also present in the test set and there are mislabeled examples ?</p>\n<p>Did it helped others to handle the class imbalance problem ? If yes then how much ?</p>",
      "rawMarkdown": "Hello everyone I'm a beginner, I was doing this challenge recently and I tried to make a simple approach and I found the following results :-\n\n1. **baseline**(ResNet50_256_image_size_30_epochs): (public/private) (0.7605/0.7483)\n1. baseline+**class_weights**(class weights for class imbalance): doesn't helps (0.7375/0.7542)\n1. baseline+**512_image_size**: does helps (public/private) (0.7875/0.7938)\n1. baseline+**label_smoothing**(label smoothing handling training with noisy data): does helps too (public/private) (0.7718/0.7748)\n1. baseline+label_smoothing+**data_augmentation+dropout+60_epochs**: does helps too (0.8253/0.8383)\n\n**Note that in all these experiments I saw overfitting .**\n\nAs you can see that using class_weights for handling class imbalance helped with private score on LB but it performed worse on the public score on LB.\n\nI read other's results and by what I understood was that they improved their results using center_cropping, TTA, using data augmentations, ensembling , but only a few people tried to focus on the class imbalance problem , is it not important in this case ? probably because of the class imbalance is also present in the test set and there are mislabeled examples ?\n\nDid it helped others to handle the class imbalance problem ? If yes then how much ?",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2766248": "Hello everyone I'm a beginner, I was doing this challenge recently and I tried to make a simple approach and I found the following results :-\n\n1. **baseline**(ResNet50_256_image_size_30_epochs): (public/private) (0.7605/0.7483)\n1. baseline+**class_weights**(class weights for class imbalance): doesn't helps (0.7375/0.7542)\n1. baseline+**512_image_size**: does helps (public/private) (0.7875/0.7938)\n1. baseline+**label_smoothing**(label smoothing handling training with noisy data): does helps too (public/private) (0.7718/0.7748)\n1. baseline+label_smoothing+**data_augmentation+dropout+60_epochs**: does helps too (0.8253/0.8383)\n\n**Note that in all these experiments I saw overfitting .**\n\nAs you can see that using class_weights for handling class imbalance helped with private score on LB but it performed worse on the public score on LB.\n\nI read other's results and by what I understood was that they improved their results using center_cropping, TTA, using data augmentations, ensembling , but only a few people tried to focus on the class imbalance problem , is it not important in this case ? probably because of the class imbalance is also present in the test set and there are mislabeled examples ?\n\nDid it helped others to handle the class imbalance problem ? If yes then how much ?"
  }
}