{
  "id": 316249,
  "title": "How to properly address class imbalance? Class reweighting techinques for training long-tailed data.",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/316249",
  "author_name": "John Park",
  "post_date": "2022-04-01T01:58:22.793000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p><strong>I would like to share a paper that addresses the long-tail class distribution problem.</strong></p>\n<p>Before I begin, I want to point out that our 2022 dataset is less skewed than our previous datasets, such as 2021, in which the number of training images ranges from 1 to 3,000. For Herbarium 2022, it ranges from 5 to 80, which is much better to handle. We made this effort in the hopes of prototyping \"specialist\" models that are super reliable for inferencing plants in specific geographical regions, which can actually help botanists in herbarium in many ways.  </p>\n<p>Nonetheless, our 2022 data is still imbalanced, and I think it is quite important to address this in the training and testing pipeline. </p>\n<p><strong>The paper I will be sharing today</strong>, <a href=\"https://arxiv.org/abs/1901.05555\" target=\"_blank\">Class-Balanced Loss Based on Effective Number of Samples, by Yin Chui et al (2019)</a>, <strong>presents a simple class reweighting technique that works well for many imbalanced fine-grained data, such as long-tailed CIFAR, iNaturalist, and ImageNet.</strong></p>\n<p>The author briefly talks about two different methods to address the class imbalance, sampling and reweighting. In their opinion, none of them so far work best, especially with larger datasets with high number of classes. For sampling, the under-sampling strategy increases the chance of skipping important data, and over-sampling is prone to overfitting. <strong>For reweighting, the classic inverse class frequency weight actually yields worse performance in many cases.</strong></p>\n<p><strong>They focus on the speculation that perhaps the inverse class-frequency method weighs the less represented classes much heavier than necessary.</strong> For example, in the medical field, ttaking a square root of inverse class frequencies appears to be working better, by smoothing the difference in weights between the less represented and the more represented classes.  </p>\n<p>The authors take this problem further to develop an idea for the \"effective number of samples\", and build a quite profound theoretical framework. <strong>The key point is that instead of taking the face value from the sampling class frequency, considering a new number of samples that really matter for truly differentiating different classes.</strong> And then, they smooth the weight distribution based on the \"effective number of samples\". <strong>By doing this, you are essentially finding a new optimal weight distribution between the no-weight and inverse class-frequency weight, to re-weight the classes.</strong></p>\n<p>Now I should really type in some equations here but I am not super familiar of the discussion settings here so let me do this briefly:</p>\n<blockquote>\n  <ul>\n  <li>Class frequency<br>\n  cls_freq : class frequency of given class </li>\n  <li>Effective number EffN<br>\n  EffN= (1-beta**cls_freq)/(1-beta)</li>\n  <li>Hyperparameter beta<br>\n  beta=(N-1)/N</li>\n  <li>The weights for each class<br>\n  w = 1/EffN</li>\n  </ul>\n</blockquote>\n<p>Here N is the realized number of classes and the hyperparameter to decide the degree of weight smoothing.  When N=1, the weight is set to 1, which is the normal case. As N goes to infinity, the effective number equals the class frequency and the weights are exactly the same as the inverse class-frequency of the classes. <strong>You can find a middle ground between the no-weight and inverse class-frequency weight by adjusting N (so beta).</strong></p>\n<p>The authors used <strong>weighted cross-entropy loss</strong> and <strong>weighted focal loss</strong> for their experiments <strong>by simply re-weighting the sample loss by replacing the weight with inverse of effective numbers</strong>. They demonstrate this simple and less severe new re-weighting method improves the performance of neural networks on class imbalanced data.</p>\n<p>I think the highlight of this paper is when they showed the performance of their method on iNaturlist and imageNet both beating the regular CE loss. <strong>Their weighted focal loss with gamma=0.5 and beta=0.999 worked the best with the standard training method</strong> they adapted from <a href=\"https://arxiv.org/abs/1706.02677\" target=\"_blank\">Goyal et al. (2017) paper from facebook: Accurate, large minibatch sgd: training imagenet in 1 hour</a>. <strong>The margin was about 3-4%</strong>, which I think could make a lot of difference based on the current leaderboard. </p>\n<p>That's it! I think the paper is very well written and has lots of useful information for the field of study. Highly recommend it to anyone who's interested to explore more. It also shows some cases in which high N (i.e. beta=0.99 and 0.999) shows worse performance, and low N (beta =0.9) shows better performance, which resembles the worse performance of the inverse class-frequency weights. <strong>The authors argue that a fine-grained dataset probably has less number of unique prototypes thus requires more delicate treatment for the weights.</strong> I think it is an interesting point. Since our dataset is exclusively about plant specimens, it would be super interesting to see what the optimal beta is for our data and to see if that well corresponds to the theoretical framework that the authors provided.</p>\n<p>That is really it! I hope this helps. I am planning on coming up with a notebook that applies this simple weighted loss in the training loop, probably using TensorFlow 2. Stay tuned and happy modeling! </p>",
  "messages": [
    {
      "id": 1741602,
      "postDate": "2022-04-01T01:58:22.793Z",
      "content": "<p>Hi everyone, </p>\n<p><strong>I would like to share a paper that addresses the long-tail class distribution problem.</strong></p>\n<p>Before I begin, I want to point out that our 2022 dataset is less skewed than our previous datasets, such as 2021, in which the number of training images ranges from 1 to 3,000. For Herbarium 2022, it ranges from 5 to 80, which is much better to handle. We made this effort in the hopes of prototyping \"specialist\" models that are super reliable for inferencing plants in specific geographical regions, which can actually help botanists in herbarium in many ways.  </p>\n<p>Nonetheless, our 2022 data is still imbalanced, and I think it is quite important to address this in the training and testing pipeline. </p>\n<p><strong>The paper I will be sharing today</strong>, <a href=\"https://arxiv.org/abs/1901.05555\" target=\"_blank\">Class-Balanced Loss Based on Effective Number of Samples, by Yin Chui et al (2019)</a>, <strong>presents a simple class reweighting technique that works well for many imbalanced fine-grained data, such as long-tailed CIFAR, iNaturalist, and ImageNet.</strong></p>\n<p>The author briefly talks about two different methods to address the class imbalance, sampling and reweighting. In their opinion, none of them so far work best, especially with larger datasets with high number of classes. For sampling, the under-sampling strategy increases the chance of skipping important data, and over-sampling is prone to overfitting. <strong>For reweighting, the classic inverse class frequency weight actually yields worse performance in many cases.</strong></p>\n<p><strong>They focus on the speculation that perhaps the inverse class-frequency method weighs the less represented classes much heavier than necessary.</strong> For example, in the medical field, ttaking a square root of inverse class frequencies appears to be working better, by smoothing the difference in weights between the less represented and the more represented classes.  </p>\n<p>The authors take this problem further to develop an idea for the \"effective number of samples\", and build a quite profound theoretical framework. <strong>The key point is that instead of taking the face value from the sampling class frequency, considering a new number of samples that really matter for truly differentiating different classes.</strong> And then, they smooth the weight distribution based on the \"effective number of samples\". <strong>By doing this, you are essentially finding a new optimal weight distribution between the no-weight and inverse class-frequency weight, to re-weight the classes.</strong></p>\n<p>Now I should really type in some equations here but I am not super familiar of the discussion settings here so let me do this briefly:</p>\n<blockquote>\n  <ul>\n  <li>Class frequency<br>\n  cls_freq : class frequency of given class </li>\n  <li>Effective number EffN<br>\n  EffN= (1-beta**cls_freq)/(1-beta)</li>\n  <li>Hyperparameter beta<br>\n  beta=(N-1)/N</li>\n  <li>The weights for each class<br>\n  w = 1/EffN</li>\n  </ul>\n</blockquote>\n<p>Here N is the realized number of classes and the hyperparameter to decide the degree of weight smoothing.  When N=1, the weight is set to 1, which is the normal case. As N goes to infinity, the effective number equals the class frequency and the weights are exactly the same as the inverse class-frequency of the classes. <strong>You can find a middle ground between the no-weight and inverse class-frequency weight by adjusting N (so beta).</strong></p>\n<p>The authors used <strong>weighted cross-entropy loss</strong> and <strong>weighted focal loss</strong> for their experiments <strong>by simply re-weighting the sample loss by replacing the weight with inverse of effective numbers</strong>. They demonstrate this simple and less severe new re-weighting method improves the performance of neural networks on class imbalanced data.</p>\n<p>I think the highlight of this paper is when they showed the performance of their method on iNaturlist and imageNet both beating the regular CE loss. <strong>Their weighted focal loss with gamma=0.5 and beta=0.999 worked the best with the standard training method</strong> they adapted from <a href=\"https://arxiv.org/abs/1706.02677\" target=\"_blank\">Goyal et al. (2017) paper from facebook: Accurate, large minibatch sgd: training imagenet in 1 hour</a>. <strong>The margin was about 3-4%</strong>, which I think could make a lot of difference based on the current leaderboard. </p>\n<p>That's it! I think the paper is very well written and has lots of useful information for the field of study. Highly recommend it to anyone who's interested to explore more. It also shows some cases in which high N (i.e. beta=0.99 and 0.999) shows worse performance, and low N (beta =0.9) shows better performance, which resembles the worse performance of the inverse class-frequency weights. <strong>The authors argue that a fine-grained dataset probably has less number of unique prototypes thus requires more delicate treatment for the weights.</strong> I think it is an interesting point. Since our dataset is exclusively about plant specimens, it would be super interesting to see what the optimal beta is for our data and to see if that well corresponds to the theoretical framework that the authors provided.</p>\n<p>That is really it! I hope this helps. I am planning on coming up with a notebook that applies this simple weighted loss in the training loop, probably using TensorFlow 2. Stay tuned and happy modeling! </p>",
      "rawMarkdown": "Hi everyone, \n\n**I would like to share a paper that addresses the long-tail class distribution problem.**\n\nBefore I begin, I want to point out that our 2022 dataset is less skewed than our previous datasets, such as 2021, in which the number of training images ranges from 1 to 3,000. For Herbarium 2022, it ranges from 5 to 80, which is much better to handle. We made this effort in the hopes of prototyping \"specialist\" models that are super reliable for inferencing plants in specific geographical regions, which can actually help botanists in herbarium in many ways.  \n\nNonetheless, our 2022 data is still imbalanced, and I think it is quite important to address this in the training and testing pipeline. \n\n**The paper I will be sharing today**, [Class-Balanced Loss Based on Effective Number of Samples, by Yin Chui et al (2019)](https://arxiv.org/abs/1901.05555), **presents a simple class reweighting technique that works well for many imbalanced fine-grained data, such as long-tailed CIFAR, iNaturalist, and ImageNet.**\n\nThe author briefly talks about two different methods to address the class imbalance, sampling and reweighting. In their opinion, none of them so far work best, especially with larger datasets with high number of classes. For sampling, the under-sampling strategy increases the chance of skipping important data, and over-sampling is prone to overfitting. **For reweighting, the classic inverse class frequency weight actually yields worse performance in many cases.**\n\n**They focus on the speculation that perhaps the inverse class-frequency method weighs the less represented classes much heavier than necessary.** For example, in the medical field, ttaking a square root of inverse class frequencies appears to be working better, by smoothing the difference in weights between the less represented and the more represented classes.  \n\nThe authors take this problem further to develop an idea for the \"effective number of samples\", and build a quite profound theoretical framework. **The key point is that instead of taking the face value from the sampling class frequency, considering a new number of samples that really matter for truly differentiating different classes.** And then, they smooth the weight distribution based on the \"effective number of samples\". **By doing this, you are essentially finding a new optimal weight distribution between the no-weight and inverse class-frequency weight, to re-weight the classes.**\n\n\nNow I should really type in some equations here but I am not super familiar of the discussion settings here so let me do this briefly:\n\n>-  Class frequency\n> cls_freq : class frequency of given class \n> - Effective number EffN\n> EffN= (1-beta**cls_freq)/(1-beta)\n> - Hyperparameter beta\n> beta=(N-1)/N\n> - The weights for each class\n> w = 1/EffN\n\nHere N is the realized number of classes and the hyperparameter to decide the degree of weight smoothing.  When N=1, the weight is set to 1, which is the normal case. As N goes to infinity, the effective number equals the class frequency and the weights are exactly the same as the inverse class-frequency of the classes. **You can find a middle ground between the no-weight and inverse class-frequency weight by adjusting N (so beta).**\n\nThe authors used **weighted cross-entropy loss** and **weighted focal loss** for their experiments **by simply re-weighting the sample loss by replacing the weight with inverse of effective numbers**. They demonstrate this simple and less severe new re-weighting method improves the performance of neural networks on class imbalanced data.\n\nI think the highlight of this paper is when they showed the performance of their method on iNaturlist and imageNet both beating the regular CE loss. **Their weighted focal loss with gamma=0.5 and beta=0.999 worked the best with the standard training method** they adapted from [Goyal et al. (2017) paper from facebook: Accurate, large minibatch sgd: training imagenet in 1 hour](https://arxiv.org/abs/1706.02677). **The margin was about 3-4%**, which I think could make a lot of difference based on the current leaderboard. \n\nThat's it! I think the paper is very well written and has lots of useful information for the field of study. Highly recommend it to anyone who's interested to explore more. It also shows some cases in which high N (i.e. beta=0.99 and 0.999) shows worse performance, and low N (beta =0.9) shows better performance, which resembles the worse performance of the inverse class-frequency weights. **The authors argue that a fine-grained dataset probably has less number of unique prototypes thus requires more delicate treatment for the weights.** I think it is an interesting point. Since our dataset is exclusively about plant specimens, it would be super interesting to see what the optimal beta is for our data and to see if that well corresponds to the theoretical framework that the authors provided.\n\nThat is really it! I hope this helps. I am planning on coming up with a notebook that applies this simple weighted loss in the training loop, probably using TensorFlow 2. Stay tuned and happy modeling! \n\n\n\n\n\n\n\n   "
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1741602": "Hi everyone, \n\n**I would like to share a paper that addresses the long-tail class distribution problem.**\n\nBefore I begin, I want to point out that our 2022 dataset is less skewed than our previous datasets, such as 2021, in which the number of training images ranges from 1 to 3,000. For Herbarium 2022, it ranges from 5 to 80, which is much better to handle. We made this effort in the hopes of prototyping \"specialist\" models that are super reliable for inferencing plants in specific geographical regions, which can actually help botanists in herbarium in many ways.  \n\nNonetheless, our 2022 data is still imbalanced, and I think it is quite important to address this in the training and testing pipeline. \n\n**The paper I will be sharing today**, [Class-Balanced Loss Based on Effective Number of Samples, by Yin Chui et al (2019)](https://arxiv.org/abs/1901.05555), **presents a simple class reweighting technique that works well for many imbalanced fine-grained data, such as long-tailed CIFAR, iNaturalist, and ImageNet.**\n\nThe author briefly talks about two different methods to address the class imbalance, sampling and reweighting. In their opinion, none of them so far work best, especially with larger datasets with high number of classes. For sampling, the under-sampling strategy increases the chance of skipping important data, and over-sampling is prone to overfitting. **For reweighting, the classic inverse class frequency weight actually yields worse performance in many cases.**\n\n**They focus on the speculation that perhaps the inverse class-frequency method weighs the less represented classes much heavier than necessary.** For example, in the medical field, ttaking a square root of inverse class frequencies appears to be working better, by smoothing the difference in weights between the less represented and the more represented classes.  \n\nThe authors take this problem further to develop an idea for the \"effective number of samples\", and build a quite profound theoretical framework. **The key point is that instead of taking the face value from the sampling class frequency, considering a new number of samples that really matter for truly differentiating different classes.** And then, they smooth the weight distribution based on the \"effective number of samples\". **By doing this, you are essentially finding a new optimal weight distribution between the no-weight and inverse class-frequency weight, to re-weight the classes.**\n\n\nNow I should really type in some equations here but I am not super familiar of the discussion settings here so let me do this briefly:\n\n>-  Class frequency\n> cls_freq : class frequency of given class \n> - Effective number EffN\n> EffN= (1-beta**cls_freq)/(1-beta)\n> - Hyperparameter beta\n> beta=(N-1)/N\n> - The weights for each class\n> w = 1/EffN\n\nHere N is the realized number of classes and the hyperparameter to decide the degree of weight smoothing.  When N=1, the weight is set to 1, which is the normal case. As N goes to infinity, the effective number equals the class frequency and the weights are exactly the same as the inverse class-frequency of the classes. **You can find a middle ground between the no-weight and inverse class-frequency weight by adjusting N (so beta).**\n\nThe authors used **weighted cross-entropy loss** and **weighted focal loss** for their experiments **by simply re-weighting the sample loss by replacing the weight with inverse of effective numbers**. They demonstrate this simple and less severe new re-weighting method improves the performance of neural networks on class imbalanced data.\n\nI think the highlight of this paper is when they showed the performance of their method on iNaturlist and imageNet both beating the regular CE loss. **Their weighted focal loss with gamma=0.5 and beta=0.999 worked the best with the standard training method** they adapted from [Goyal et al. (2017) paper from facebook: Accurate, large minibatch sgd: training imagenet in 1 hour](https://arxiv.org/abs/1706.02677). **The margin was about 3-4%**, which I think could make a lot of difference based on the current leaderboard. \n\nThat's it! I think the paper is very well written and has lots of useful information for the field of study. Highly recommend it to anyone who's interested to explore more. It also shows some cases in which high N (i.e. beta=0.99 and 0.999) shows worse performance, and low N (beta =0.9) shows better performance, which resembles the worse performance of the inverse class-frequency weights. **The authors argue that a fine-grained dataset probably has less number of unique prototypes thus requires more delicate treatment for the weights.** I think it is an interesting point. Since our dataset is exclusively about plant specimens, it would be super interesting to see what the optimal beta is for our data and to see if that well corresponds to the theoretical framework that the authors provided.\n\nThat is really it! I hope this helps. I am planning on coming up with a notebook that applies this simple weighted loss in the training loop, probably using TensorFlow 2. Stay tuned and happy modeling! \n\n\n\n\n\n\n\n   "
  }
}