{
  "id": 42454,
  "title": "Different ways to improve the class imbalance problem",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/42454",
  "author_name": "",
  "post_date": "2017-10-31T08:01:31.731132800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As we see , the data provided by the competition is imbalanced.</p>\n\n<p>I split my data into train(about 12,200,000 images) and val(about 10,000 images), and I ensure all data have 5270 classes. Then I compute the number of images in different categories in the train data. As shown in the picture.</p>\n\n<p>In the picture,  the min category has only 11 images , but the max category has 79698 images. And there are only about 230 categories which have more 10,000 images. That is to say, we have extreme ratio of imbalance and large portion of classes are minority. </p>\n\n<p>In this way , undersampling can perform on a par with oversampling.</p>\n\n<p>It seems that we should improve the problem of class imbalance to get better score. Then I find a \n <a href=\"http://arxiv.org/pdf/1710.05381.pdf\">paper</a> that does a  systematic study of the class imbalance problem in convolutional neural networks. The paper compares different ways to improve the class imbalance,  and drew a conclusion </p>\n\n<p>&gt; To achieve the best accuracy, one should apply thresholding to compensate for prior class probabilities. A combination of thresholding with baseline and oversampling is the most preferable, whereas it should not be combined with undersampling.</p>\n\n<p>But this is only my find, hoping anyone else to show more and  better ways to improve the class imbalance problem.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "237893",
      "postDate": "10/31/2017 08:01:31",
      "content": "<p>As we see , the data provided by the competition is imbalanced.</p>\n\n<p>I split my data into train(about 12,200,000 images) and val(about 10,000 images), and I ensure all data have 5270 classes. Then I compute the number of images in different categories in the train data. As shown in the picture.</p>\n\n<p>In the picture,  the min category has only 11 images , but the max category has 79698 images. And there are only about 230 categories which have more 10,000 images. That is to say, we have extreme ratio of imbalance and large portion of classes are minority. </p>\n\n<p>In this way , undersampling can perform on a par with oversampling.</p>\n\n<p>It seems that we should improve the problem of class imbalance to get better score. Then I find a \n <a href=\"http://arxiv.org/pdf/1710.05381.pdf\">paper</a> that does a  systematic study of the class imbalance problem in convolutional neural networks. The paper compares different ways to improve the class imbalance,  and drew a conclusion </p>\n\n<p>&gt; To achieve the best accuracy, one should apply thresholding to compensate for prior class probabilities. A combination of thresholding with baseline and oversampling is the most preferable, whereas it should not be combined with undersampling.</p>\n\n<p>But this is only my find, hoping anyone else to show more and  better ways to improve the class imbalance problem.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "As we see , the data provided by the competition is imbalanced.\n\nI split my data into train(about 12,200,000 images) and val(about 10,000 images), and I ensure all data have 5270 classes. Then I compute the number of images in different categories in the train data. As shown in the picture.\n\nIn the picture,  the min category has only 11 images , but the max category has 79698 images. And there are only about 230 categories which have more 10,000 images. That is to say, we have extreme ratio of imbalance and large portion of classes are minority. \n\nIn this way , undersampling can perform on a par with oversampling.\n\nIt seems that we should improve the problem of class imbalance to get better score. Then I find a \n [paper][1] that does a  systematic study of the class imbalance problem in convolutional neural networks. The paper compares different ways to improve the class imbalance,  and drew a conclusion \n\n&gt; To achieve the best accuracy, one should apply thresholding to compensate for prior class probabilities. A combination of thresholding with baseline and oversampling is the most preferable, whereas it should not be combined with undersampling.\n\nBut this is only my find, hoping anyone else to show more and  better ways to improve the class imbalance problem.\n\nThanks!\n\n\n  [1]: http://arxiv.org/pdf/1710.05381.pdf",
      "votes": null
    },
    {
      "id": "237907",
      "postDate": "10/31/2017 09:34:14",
      "content": "<p>I recently plotted the precision/recall for my best model. (I did not do anything special to deal with the class imbalance when I trained this model.) It shows that quite a few of the small categories in my validation set have a perfect prediction score, but also quite a few have zero correct predictions. And of those with zero correct predictions, the categories it predicted actually did make a lot of sense. So some small categories work just fine, while others don't.</p>\n\n<p>For example, \"TELEVISION LCD\" had zero correct predictions because pretty much all the images in this category were predicted to be \"TELEVISION LED\" -- now, I can't look at just a picture of a TV and tell you whether it's an LCD or LED TV, so no wonder the model can't either. </p>\n\n<p>None of the predictions for the small categories were \"obviously\" wrong and often they were quite reasonable. It's just that they did not match the exact category, but often they matched one very close to it.</p>\n\n<p>By looking at what my model actually predicted and the images it had to work with, I now believe that the class imbalance is not really as big a problem as I thought it was. The bigger problem is that the categories are often confusing, the images are often not very good or ambiguous (true, this is a bigger problem for smaller categories than for big categories), and certain categories are just very hard (music and books).</p>",
      "rawMarkdown": "I recently plotted the precision/recall for my best model. (I did not do anything special to deal with the class imbalance when I trained this model.) It shows that quite a few of the small categories in my validation set have a perfect prediction score, but also quite a few have zero correct predictions. And of those with zero correct predictions, the categories it predicted actually did make a lot of sense. So some small categories work just fine, while others don't.\n\nFor example, \"TELEVISION LCD\" had zero correct predictions because pretty much all the images in this category were predicted to be \"TELEVISION LED\" -- now, I can't look at just a picture of a TV and tell you whether it's an LCD or LED TV, so no wonder the model can't either. \n\nNone of the predictions for the small categories were \"obviously\" wrong and often they were quite reasonable. It's just that they did not match the exact category, but often they matched one very close to it.\n\nBy looking at what my model actually predicted and the images it had to work with, I now believe that the class imbalance is not really as big a problem as I thought it was. The bigger problem is that the categories are often confusing, the images are often not very good or ambiguous (true, this is a bigger problem for smaller categories than for big categories), and certain categories are just very hard (music and books).",
      "votes": null
    },
    {
      "id": "237919",
      "postDate": "10/31/2017 10:07:43",
      "content": "<p>I completely agree with you @human_Analog\nThis is also the case for ''non rare'' categories. For instance, how would you discriminate between books genre or music genre (e.g.,  {CD HARD ROCK - CD METAL} Vs {CD SOUL - CD FUNK - CD DISCO}). There is no other way to extract meaningful features (e.g., text) as the resolution is too low.\nPlus I saw multiple examples that were ''wrongly'' labeled (Bag Vs Backpack)</p>",
      "rawMarkdown": "I completely agree with you @human_Analog\nThis is also the case for ''non rare'' categories. For instance, how would you discriminate between books genre or music genre (e.g.,  {CD HARD ROCK - CD METAL} Vs {CD SOUL - CD FUNK - CD DISCO}). There is no other way to extract meaningful features (e.g., text) as the resolution is too low.\nPlus I saw multiple examples that were ''wrongly'' labeled (Bag Vs Backpack)",
      "votes": null
    },
    {
      "id": "501258",
      "postDate": "03/27/2019 05:35:48",
      "content": "<p>what about focal loss have u tried that approach to solve problem it works really well for hard examples  to classify , because it focuses on minorities than majority classes</p>",
      "rawMarkdown": "what about focal loss have u tried that approach to solve problem it works really well for hard examples  to classify , because it focuses on minorities than majority classes",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 237907,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "10/31/2017 09:34:14",
      "content": "<p>I recently plotted the precision/recall for my best model. (I did not do anything special to deal with the class imbalance when I trained this model.) It shows that quite a few of the small categories in my validation set have a perfect prediction score, but also quite a few have zero correct predictions. And of those with zero correct predictions, the categories it predicted actually did make a lot of sense. So some small categories work just fine, while others don't.</p>\n\n<p>For example, \"TELEVISION LCD\" had zero correct predictions because pretty much all the images in this category were predicted to be \"TELEVISION LED\" -- now, I can't look at just a picture of a TV and tell you whether it's an LCD or LED TV, so no wonder the model can't either. </p>\n\n<p>None of the predictions for the small categories were \"obviously\" wrong and often they were quite reasonable. It's just that they did not match the exact category, but often they matched one very close to it.</p>\n\n<p>By looking at what my model actually predicted and the images it had to work with, I now believe that the class imbalance is not really as big a problem as I thought it was. The bigger problem is that the categories are often confusing, the images are often not very good or ambiguous (true, this is a bigger problem for smaller categories than for big categories), and certain categories are just very hard (music and books).</p>",
      "votes": null,
      "replies": [
        {
          "id": 237919,
          "author_name": "elberi",
          "author_url": "",
          "post_date": "10/31/2017 10:07:43",
          "content": "<p>I completely agree with you @human_Analog\nThis is also the case for ''non rare'' categories. For instance, how would you discriminate between books genre or music genre (e.g.,  {CD HARD ROCK - CD METAL} Vs {CD SOUL - CD FUNK - CD DISCO}). There is no other way to extract meaningful features (e.g., text) as the resolution is too low.\nPlus I saw multiple examples that were ''wrongly'' labeled (Bag Vs Backpack)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 501258,
      "author_name": "rohandx1996",
      "author_url": "",
      "post_date": "03/27/2019 05:35:48",
      "content": "<p>what about focal loss have u tried that approach to solve problem it works really well for hard examples  to classify , because it focuses on minorities than majority classes</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "237893": "As we see , the data provided by the competition is imbalanced.\n\nI split my data into train(about 12,200,000 images) and val(about 10,000 images), and I ensure all data have 5270 classes. Then I compute the number of images in different categories in the train data. As shown in the picture.\n\nIn the picture,  the min category has only 11 images , but the max category has 79698 images. And there are only about 230 categories which have more 10,000 images. That is to say, we have extreme ratio of imbalance and large portion of classes are minority. \n\nIn this way , undersampling can perform on a par with oversampling.\n\nIt seems that we should improve the problem of class imbalance to get better score. Then I find a \n [paper][1] that does a  systematic study of the class imbalance problem in convolutional neural networks. The paper compares different ways to improve the class imbalance,  and drew a conclusion \n\n&gt; To achieve the best accuracy, one should apply thresholding to compensate for prior class probabilities. A combination of thresholding with baseline and oversampling is the most preferable, whereas it should not be combined with undersampling.\n\nBut this is only my find, hoping anyone else to show more and  better ways to improve the class imbalance problem.\n\nThanks!\n\n\n  [1]: http://arxiv.org/pdf/1710.05381.pdf",
    "237907": "I recently plotted the precision/recall for my best model. (I did not do anything special to deal with the class imbalance when I trained this model.) It shows that quite a few of the small categories in my validation set have a perfect prediction score, but also quite a few have zero correct predictions. And of those with zero correct predictions, the categories it predicted actually did make a lot of sense. So some small categories work just fine, while others don't.\n\nFor example, \"TELEVISION LCD\" had zero correct predictions because pretty much all the images in this category were predicted to be \"TELEVISION LED\" -- now, I can't look at just a picture of a TV and tell you whether it's an LCD or LED TV, so no wonder the model can't either. \n\nNone of the predictions for the small categories were \"obviously\" wrong and often they were quite reasonable. It's just that they did not match the exact category, but often they matched one very close to it.\n\nBy looking at what my model actually predicted and the images it had to work with, I now believe that the class imbalance is not really as big a problem as I thought it was. The bigger problem is that the categories are often confusing, the images are often not very good or ambiguous (true, this is a bigger problem for smaller categories than for big categories), and certain categories are just very hard (music and books).",
    "237919": "I completely agree with you @human_Analog\nThis is also the case for ''non rare'' categories. For instance, how would you discriminate between books genre or music genre (e.g.,  {CD HARD ROCK - CD METAL} Vs {CD SOUL - CD FUNK - CD DISCO}). There is no other way to extract meaningful features (e.g., text) as the resolution is too low.\nPlus I saw multiple examples that were ''wrongly'' labeled (Bag Vs Backpack)",
    "501258": "what about focal loss have u tried that approach to solve problem it works really well for hard examples  to classify , because it focuses on minorities than majority classes"
  },
  "source": "meta"
}