{
  "id": 172892,
  "title": "Why do people care about balancing classes?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172892",
  "author_name": "CPMP",
  "post_date": "2020-08-06T21:44:36.715000",
  "votes": 52,
  "comment_count": 77,
  "views": 0,
  "content": "<p>It is a genuine question for me.  I have tried a bit upsampling or weighting, and did not get real improvement.</p>\n<p>Why are people so obsessed with the class imbalance?  </p>\n<p>I would understand if we care about calibrating predictions.  But we don't. Indeed, roc-auc only depend son the ordering of predictions, not their values.</p>\n<p>the only reason I see is to add diversity in models, which is good when ensembling.  But there are many ways to add diversity, and they are way less discussed than class imbalance.</p>\n<p>I'll try again when I will be using larger images and larger models, but so far it looks like focusing on the wrong issue to me.  </p>\n<p><strong>Edit.</strong>  Most comments are very interesting and worth reading.  There is some confusion about wording however.  Upsampling the minority class means that samples form minority class are sampled more than once (duplicated).  Obviously. some confuse that with adding new samples for the minority class. Wording isn't really an issue as long as we discuss the same topic.</p>\n<p>Adding more data helps in general. Adding only samples from minority class does not help necessarily however.  I have experienced, and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> too, that only adding positive samples can hurt actually. </p>\n<p>Anyway, let's leave it at this:  adding external data helps in general. </p>\n<p>This post was not about adding external data.  It was about usampling minority class, i.e. stick with the competition data but use Melanoma samples more than the others when training.  I have not found it to be useful here, and my experience is that it is never useful when using deep learning.  I did found it useful with, say XGBoost, in the past.</p>\n<p>There is one reason for upsampling that is worth discussing IMHO, see comments from <a href=\"https://www.kaggle.com/Optimo\" target=\"_blank\">@Optimo</a>: if batch size is small enough, then many batches have only samples from the majority class.  Some loss functions don't even work in that case, esp the ones approximating auc directly.  My take here is that if you use a loss that works in that case, eg BCE, then this averages out over many batches.  And, as was also pointed to in comments, one can use gradient accumulation to make sure logical batches are large enough.</p>\n<p><strong>Second edit.</strong>  It seems the main benefit, according to several comments, is that training converges faster when upsampling the minority class.  This alone is a good reason for upsampling.  </p>\n<p>Unless mistaken no one claimed they got better roc-auc with upsampling for CNN models.  </p>\n<p><strong>Third edit.</strong> Chris Deotte says he ran experiment showing a 0.002 CV AUC improvement with upsampling.  Looking forward to his experiment details.  Anyway, this is the first time someone claims that upsampling improved roc-auc for CNN.  I'll try harder then given I did not found it useful till now.  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130#964575\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130#964575</a></p>\n<p><strong>Fourth edit.</strong>  <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> commented that oversampling the minority class can improve mixup, see his comment below.  </p>",
  "messages": [
    {
      "id": 961043,
      "postDate": "2020-08-06T21:44:36.717Z",
      "content": "<p>It is a genuine question for me.  I have tried a bit upsampling or weighting, and did not get real improvement.</p>\n<p>Why are people so obsessed with the class imbalance?  </p>\n<p>I would understand if we care about calibrating predictions.  But we don't. Indeed, roc-auc only depend son the ordering of predictions, not their values.</p>\n<p>the only reason I see is to add diversity in models, which is good when ensembling.  But there are many ways to add diversity, and they are way less discussed than class imbalance.</p>\n<p>I'll try again when I will be using larger images and larger models, but so far it looks like focusing on the wrong issue to me.  </p>\n<p><strong>Edit.</strong>  Most comments are very interesting and worth reading.  There is some confusion about wording however.  Upsampling the minority class means that samples form minority class are sampled more than once (duplicated).  Obviously. some confuse that with adding new samples for the minority class. Wording isn't really an issue as long as we discuss the same topic.</p>\n<p>Adding more data helps in general. Adding only samples from minority class does not help necessarily however.  I have experienced, and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> too, that only adding positive samples can hurt actually. </p>\n<p>Anyway, let's leave it at this:  adding external data helps in general. </p>\n<p>This post was not about adding external data.  It was about usampling minority class, i.e. stick with the competition data but use Melanoma samples more than the others when training.  I have not found it to be useful here, and my experience is that it is never useful when using deep learning.  I did found it useful with, say XGBoost, in the past.</p>\n<p>There is one reason for upsampling that is worth discussing IMHO, see comments from <a href=\"https://www.kaggle.com/Optimo\" target=\"_blank\">@Optimo</a>: if batch size is small enough, then many batches have only samples from the majority class.  Some loss functions don't even work in that case, esp the ones approximating auc directly.  My take here is that if you use a loss that works in that case, eg BCE, then this averages out over many batches.  And, as was also pointed to in comments, one can use gradient accumulation to make sure logical batches are large enough.</p>\n<p><strong>Second edit.</strong>  It seems the main benefit, according to several comments, is that training converges faster when upsampling the minority class.  This alone is a good reason for upsampling.  </p>\n<p>Unless mistaken no one claimed they got better roc-auc with upsampling for CNN models.  </p>\n<p><strong>Third edit.</strong> Chris Deotte says he ran experiment showing a 0.002 CV AUC improvement with upsampling.  Looking forward to his experiment details.  Anyway, this is the first time someone claims that upsampling improved roc-auc for CNN.  I'll try harder then given I did not found it useful till now.  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130#964575\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130#964575</a></p>\n<p><strong>Fourth edit.</strong>  <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> commented that oversampling the minority class can improve mixup, see his comment below.  </p>",
      "rawMarkdown": "It is a genuine question for me.  I have tried a bit upsampling or weighting, and did not get real improvement.\n\nWhy are people so obsessed with the class imbalance?  \n\nI would understand if we care about calibrating predictions.  But we don't. Indeed, roc-auc only depend son the ordering of predictions, not their values.\n\nthe only reason I see is to add diversity in models, which is good when ensembling.  But there are many ways to add diversity, and they are way less discussed than class imbalance.\n\nI'll try again when I will be using larger images and larger models, but so far it looks like focusing on the wrong issue to me.  \n\n**Edit.**  Most comments are very interesting and worth reading.  There is some confusion about wording however.  Upsampling the minority class means that samples form minority class are sampled more than once (duplicated).  Obviously. some confuse that with adding new samples for the minority class. Wording isn't really an issue as long as we discuss the same topic.\n\nAdding more data helps in general. Adding only samples from minority class does not help necessarily however.  I have experienced, and @philippsinger too, that only adding positive samples can hurt actually. \n\nAnyway, let's leave it at this:  adding external data helps in general. \n\nThis post was not about adding external data.  It was about usampling minority class, i.e. stick with the competition data but use Melanoma samples more than the others when training.  I have not found it to be useful here, and my experience is that it is never useful when using deep learning.  I did found it useful with, say XGBoost, in the past.\n\nThere is one reason for upsampling that is worth discussing IMHO, see comments from @Optimo: if batch size is small enough, then many batches have only samples from the majority class.  Some loss functions don't even work in that case, esp the ones approximating auc directly.  My take here is that if you use a loss that works in that case, eg BCE, then this averages out over many batches.  And, as was also pointed to in comments, one can use gradient accumulation to make sure logical batches are large enough.\n\n**Second edit.**  It seems the main benefit, according to several comments, is that training converges faster when upsampling the minority class.  This alone is a good reason for upsampling.  \n\nUnless mistaken no one claimed they got better roc-auc with upsampling for CNN models.  \n\n**Third edit.** Chris Deotte says he ran experiment showing a 0.002 CV AUC improvement with upsampling.  Looking forward to his experiment details.  Anyway, this is the first time someone claims that upsampling improved roc-auc for CNN.  I'll try harder then given I did not found it useful till now.  https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130#964575\n\n**Fourth edit.**  @nyanpn commented that oversampling the minority class can improve mixup, see his comment below.  ",
      "votes": 52
    },
    {
      "id": 961492,
      "postDate": "2020-08-07T08:26:22.490Z",
      "content": "<p>Doing nothing in imbalanced problems is often the best.</p>",
      "rawMarkdown": "Doing nothing in imbalanced problems is often the best.",
      "votes": 18
    },
    {
      "id": 962051,
      "postDate": "2020-08-07T18:43:56.867Z",
      "content": "<p>Small Example I want to add. Maybe related indirectly. Noisy student paper did quite a good experiment on if it is essential to balance data or not when adding external data. In short, they :<br>\n1) Trained Effnet On Imagenet<br>\n2) Predicted Random Images from Interner(out of distribution from original Imagenet Dataset) -&gt; generate Pseudolabels<br>\n3) Added this Images and retrained further. </p>\n<p>Now when you predict Images that are not present in the imagenet dataset, you will get different class distribution. The question that was raised is essential to balance classes when retraining network? </p>\n<p>Experimental Answer: For small model -Yes! (you will gain slight boost). It was not helpful for bigger models since they have more power to learn and separate features. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2F9a949952781f1d7da2412c95cdc8faf2%2FScreen%20Shot%202020-08-07%20at%202.41.15%20PM.png?generation=1596825750152331&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2Ff96afd76d6c8301316bdae17d6f020a2%2FScreen%20Shot%202020-08-07%20at%202.41.20%20PM.png?generation=1596825764119842&amp;alt=media\" alt=\"\"></p>\n<p>As I mentioned its probably very remotely related to the topic but perhaps can give a good experimental evidence for the discussion. </p>\n<p>Paper (<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a>)</p>",
      "rawMarkdown": "Small Example I want to add. Maybe related indirectly. Noisy student paper did quite a good experiment on if it is essential to balance data or not when adding external data. In short, they :\n1) Trained Effnet On Imagenet\n2) Predicted Random Images from Interner(out of distribution from original Imagenet Dataset) -&gt; generate Pseudolabels\n3) Added this Images and retrained further. \n\nNow when you predict Images that are not present in the imagenet dataset, you will get different class distribution. The question that was raised is essential to balance classes when retraining network? \n\nExperimental Answer: For small model -Yes! (you will gain slight boost). It was not helpful for bigger models since they have more power to learn and separate features. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2F9a949952781f1d7da2412c95cdc8faf2%2FScreen%20Shot%202020-08-07%20at%202.41.15%20PM.png?generation=1596825750152331&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2Ff96afd76d6c8301316bdae17d6f020a2%2FScreen%20Shot%202020-08-07%20at%202.41.20%20PM.png?generation=1596825764119842&amp;alt=media)\n\nAs I mentioned its probably very remotely related to the topic but perhaps can give a good experimental evidence for the discussion. \n\nPaper (https://arxiv.org/pdf/1911.04252.pdf)",
      "votes": 11
    },
    {
      "id": 961490,
      "postDate": "2020-08-07T08:24:03.537Z",
      "content": "<p>I don't upsample but I do some  balancing in some of my experiments for the sole purpose to speed up training. <br>\nFor that, I discard some benign ones , but it's done online i.e when the batch is being fed to the model. Because model don't need to focus that much on benign cases and this save training time too ;)</p>",
      "rawMarkdown": "I don't upsample but I do some  balancing in some of my experiments for the sole purpose to speed up training. \nFor that, I discard some benign ones , but it's done online i.e when the batch is being fed to the model. Because model don't need to focus that much on benign cases and this save training time too ;)",
      "votes": 9,
      "replies": [
        {
          "id": 961531,
          "postDate": "2020-08-07T09:05:33.227Z",
          "content": "<p>do you discard batches with only negative samples?</p>",
          "rawMarkdown": "do you discard batches with only negative samples?"
        },
        {
          "id": 961553,
          "postDate": "2020-08-07T09:26:24.833Z",
          "content": "<p>Yep <br>\nI manage to have positive samples in every batch for those expeiments.</p>",
          "rawMarkdown": "Yep \nI manage to have positive samples in every batch for those expeiments."
        },
        {
          "id": 961560,
          "postDate": "2020-08-07T09:38:10.430Z",
          "content": "<p>cool, out of curiosity do you handle this in pytorch sampler or do you manually discard batches with only zeros?<br>\nWould you share your approach? (after competition end is fine)</p>",
          "rawMarkdown": "cool, out of curiosity do you handle this in pytorch sampler or do you manually discard batches with only zeros?\nWould you share your approach? (after competition end is fine)"
        },
        {
          "id": 961607,
          "postDate": "2020-08-07T10:30:14.033Z",
          "content": "<p>I modified a bit the Pytorch sampler in this <a href=\"https://github.com/ufoym/imbalanced-dataset-sampler\" target=\"_blank\">repo</a> ( in particular the<code>_get_label</code> method to fit the needs here )  . This works for single GPU.</p>",
          "rawMarkdown": "I modified a bit the Pytorch sampler in this [repo](https://github.com/ufoym/imbalanced-dataset-sampler) ( in particular the` _get_label` method to fit the needs here )  . This works for single GPU.",
          "votes": 2
        },
        {
          "id": 961922,
          "postDate": "2020-08-07T16:07:11.473Z",
          "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>: you may refer to this discussion regarding data sampling strategy: <a href=\"https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824\" target=\"_blank\">https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824</a></p>",
          "rawMarkdown": "@optimo: you may refer to this discussion regarding data sampling strategy: https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824",
          "votes": 1
        },
        {
          "id": 962545,
          "postDate": "2020-08-08T08:14:27.557Z",
          "content": "<p>I think this is a right way to do.</p>",
          "rawMarkdown": "I think this is a right way to do.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1368157,
      "postDate": "2021-06-28T11:14:28.087Z",
      "content": "<p>While class imbalance is not important for auc, it may make a difference for mixup, since most samples will simply be a mix of two negative samples when mixup is performed on imbalanced data.<br>\nSo I just oversampled the positive samples by x3 so that the mixup would try more combinations of positive and negative.</p>",
      "rawMarkdown": "While class imbalance is not important for auc, it may make a difference for mixup, since most samples will simply be a mix of two negative samples when mixup is performed on imbalanced data.\nSo I just oversampled the positive samples by x3 so that the mixup would try more combinations of positive and negative.",
      "votes": 3,
      "replies": [
        {
          "id": 1368246,
          "postDate": "2021-06-28T12:50:04.227Z",
          "content": "<p>I agree, and I do it now as well.  Let me edit the post.</p>",
          "rawMarkdown": "I agree, and I do it now as well.  Let me edit the post.",
          "votes": 1
        }
      ]
    },
    {
      "id": 964057,
      "postDate": "2020-08-09T14:55:51.290Z",
      "content": "<p>Imagine if there are 1000000000 negative images and just 1 positive image, then we need to do lots of epochs for back-propagating that positive image effectively.</p>",
      "rawMarkdown": "Imagine if there are 1000000000 negative images and just 1 positive image, then we need to do lots of epochs for back-propagating that positive image effectively.",
      "votes": 5,
      "replies": [
        {
          "id": 964223,
          "postDate": "2020-08-09T17:09:59.447Z",
          "content": "<p>Yes, convergence speed seems the main benefit here.  Experimenting with it.  But have you improved roc-auc with upsampling?</p>",
          "rawMarkdown": "Yes, convergence speed seems the main benefit here.  Experimenting with it.  But have you improved roc-auc with upsampling?"
        }
      ]
    },
    {
      "id": 962174,
      "postDate": "2020-08-07T21:47:23.837Z",
      "content": "<p>This seems to be something that is taught in school as massively important. Almost every recent grad I interact with believes this to be very important. My viewpoint is that maybe in some scenarios it can lead to quicker convergence but ultimately what matters is sufficient samples to learn a pattern. if you have 100 positive points and 10,000 negative points your positive points might not be enough to discover the pattern that makes those samples positives, but if you have 10,000 and 10,000,000 positive and negative points you likely have enough positive samples to discover the pattern of those points, but might struggle with convergence. The issue most often is the number of minority points, not the ratio of minority to majority at least with most modern algorithms. </p>",
      "rawMarkdown": "This seems to be something that is taught in school as massively important. Almost every recent grad I interact with believes this to be very important. My viewpoint is that maybe in some scenarios it can lead to quicker convergence but ultimately what matters is sufficient samples to learn a pattern. if you have 100 positive points and 10,000 negative points your positive points might not be enough to discover the pattern that makes those samples positives, but if you have 10,000 and 10,000,000 positive and negative points you likely have enough positive samples to discover the pattern of those points, but might struggle with convergence. The issue most often is the number of minority points, not the ratio of minority to majority at least with most modern algorithms. ",
      "votes": 5,
      "replies": [
        {
          "id": 963092,
          "postDate": "2020-08-08T16:58:09.960Z",
          "content": "<blockquote>\n  <p>This seems to be something that is taught in school as massively important. Almost every recent grad I interact with believes this to be very important.</p>\n</blockquote>\n<p>Do you think this a hangover from how logistic regression is taught?</p>",
          "rawMarkdown": "&gt; This seems to be something that is taught in school as massively important. Almost every recent grad I interact with believes this to be very important.\n\nDo you think this a hangover from how logistic regression is taught?",
          "votes": 1
        },
        {
          "id": 964425,
          "postDate": "2020-08-09T21:00:13.023Z",
          "content": "<p>That's my understanding. Traditional methods fall prey to global statistics much worse than non-parametric methods like gradient boosted decision trees. </p>",
          "rawMarkdown": "That's my understanding. Traditional methods fall prey to global statistics much worse than non-parametric methods like gradient boosted decision trees. "
        }
      ]
    },
    {
      "id": 962245,
      "postDate": "2020-08-08T00:27:49.990Z",
      "content": "<p>I ran an experiment using triple stratified CV and posted results <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\" target=\"_blank\">here</a>. Using different seeds we consistently get around 0.002 AUC increase with 25x upsample for model RAPIDS cuML kNN on image embeddings. It's small but it helps. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc79785c71cde03a928c41bd2b1cb049d%2Faucc.png?generation=1596846316898566&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I ran an experiment using triple stratified CV and posted results [here][1]. Using different seeds we consistently get around 0.002 AUC increase with 25x upsample for model RAPIDS cuML kNN on image embeddings. It's small but it helps. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc79785c71cde03a928c41bd2b1cb049d%2Faucc.png?generation=1596846316898566&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130",
      "votes": 6,
      "replies": [
        {
          "id": 962247,
          "postDate": "2020-08-08T00:33:23.503Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> just to be clear, you mean you took all malignant images (2019, new, 2020) and added them in 25 times? That's over 100,000 images right?</p>",
          "rawMarkdown": "@cdeotte just to be clear, you mean you took all malignant images (2019, new, 2020) and added them in 25 times? That's over 100,000 images right?"
        },
        {
          "id": 962251,
          "postDate": "2020-08-08T00:39:11.030Z",
          "content": "<p>No. In this plot (experiment), I only use 2020 competition data with my RAPIDS cuML kNN model. After removing duplicates, the 2020 data has 32692 images with 584 malignant for a proportion of 1.79%. Then i add 25 copies of each of 584 malignant (i.e. 14600 images). Now the training data has <code>47292 = 32692 + 25*584</code> images with 32.1% malignant.</p>",
          "rawMarkdown": "No. In this plot (experiment), I only use 2020 competition data with my RAPIDS cuML kNN model. After removing duplicates, the 2020 data has 32692 images with 584 malignant for a proportion of 1.79%. Then i add 25 copies of each of 584 malignant (i.e. 14600 images). Now the training data has `47292 = 32692 + 25*584` images with 32.1% malignant.",
          "votes": 4
        },
        {
          "id": 962269,
          "postDate": "2020-08-08T01:32:25.097Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for clarifying, and thanks for your post I saw you made on this, very informative.</p>",
          "rawMarkdown": "@cdeotte Thanks for clarifying, and thanks for your post I saw you made on this, very informative.",
          "votes": 1
        },
        {
          "id": 962758,
          "postDate": "2020-08-08T12:18:06.833Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Do I get it right that you increase AUC for knn? I am not surprised, as knn depend son the density of each class. This is an interesting find still, thanks for checking.</p>\n<p>I have yet to see an improvement of AUC for good CNN models. I'd be very happy to see someone posting an improvement.</p>\n<p>So far the only documented improvement is training convergence improved with upsampling. I didn't think of it. It is another good outcome of this discussion for me.</p>",
          "rawMarkdown": "@cdeotte Do I get it right that you increase AUC for knn? I am not surprised, as knn depend son the density of each class. This is an interesting find still, thanks for checking.\n\nI have yet to see an improvement of AUC for good CNN models. I'd be very happy to see someone posting an improvement.\n\nSo far the only documented improvement is training convergence improved with upsampling. I didn't think of it. It is another good outcome of this discussion for me."
        }
      ]
    },
    {
      "id": 961694,
      "postDate": "2020-08-07T12:25:51.210Z",
      "content": "<p>I think it depends on the problem. If you know it advance (through probing or other means) that the test dataset might have a different target distribution , then it might worth considering balancing to match that. Otherwise, most of my experiences trying to change the principal target distribution have ended up in failures. </p>\n<p>I have also seen problems when the frequency of the minority class is too small, that balancing (or up sampling that class) in the first few epochs while training NNs  can help the model converge faster.  </p>",
      "rawMarkdown": "I think it depends on the problem. If you know it advance (through probing or other means) that the test dataset might have a different target distribution , then it might worth considering balancing to match that. Otherwise, most of my experiences trying to change the principal target distribution have ended up in failures. \n\nI have also seen problems when the frequency of the minority class is too small, that balancing (or up sampling that class) in the first few epochs while training NNs  can help the model converge faster.  ",
      "votes": 4,
      "replies": [
        {
          "id": 961717,
          "postDate": "2020-08-07T12:49:58.900Z",
          "content": "<blockquote>\n  <p>test dataset might have a different target distribution , then it might worth considering balancing to match that. </p>\n</blockquote>\n<p>How is this relevant when metric is roc-auc?  I agree with you if the metric depend on probabbility calibration.</p>\n<blockquote>\n  <p>balancing (or up sampling that class) in the first few epochs while training NNs can help the model converge faster. </p>\n</blockquote>\n<p>Interesting, thanks for sharing.  I will give it a try if possiible.</p>",
          "rawMarkdown": "&gt; test dataset might have a different target distribution , then it might worth considering balancing to match that. \n\nHow is this relevant when metric is roc-auc?  I agree with you if the metric depend on probabbility calibration.\n\n&gt; balancing (or up sampling that class) in the first few epochs while training NNs can help the model converge faster. \n\nInteresting, thanks for sharing.  I will give it a try if possiible.",
          "votes": 1
        },
        {
          "id": 961720,
          "postDate": "2020-08-07T12:54:29.727Z",
          "content": "<blockquote>\n  <p>How is this relevant when metric is roc-auc? I agree with you if the metric depends on probabbility calibration.</p>\n</blockquote>\n<p>Theoretically, it is not relevant when the metric is AUC. It is more of a general observation.</p>",
          "rawMarkdown": "&gt; How is this relevant when metric is roc-auc? I agree with you if the metric depends on probabbility calibration.\n\nTheoretically, it is not relevant when the metric is AUC. It is more of a general observation.",
          "votes": 1
        },
        {
          "id": 962515,
          "postDate": "2020-08-08T07:29:00.197Z",
          "content": "<p>i can confirm the speed up in convergency. thanks for sharing the tip. </p>\n\n<p>do you think there will be benefits in continuous training with the unbalance dataset? reading the conservation above, it seems the answer is no as the metric is auc.  </p>",
          "rawMarkdown": "i can confirm the speed up in convergency. thanks for sharing the tip. \n\ndo you think there will be benefits in continuous training with the unbalance dataset? reading the conservation above, it seems the answer is no as the metric is auc.  ",
          "replies": [
            {
              "id": 962548,
              "postDate": "2020-08-08T08:16:50.450Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 961170,
      "postDate": "2020-08-07T01:00:09.237Z",
      "content": "<p>In my opinion, if the dataset has the same distribution as the real world(for example, your test set), I don't see a necessity to do sampling even though there is class imbalance.</p>",
      "rawMarkdown": "In my opinion, if the dataset has the same distribution as the real world(for example, your test set), I don't see a necessity to do sampling even though there is class imbalance.",
      "votes": 4,
      "replies": [
        {
          "id": 961455,
          "postDate": "2020-08-07T07:37:08.430Z",
          "content": "<p>That's only when AUC is the metric I hope? In a health care setting, I would care about the positive class ;)</p>",
          "rawMarkdown": "That's only when AUC is the metric I hope? In a health care setting, I would care about the positive class ;)"
        },
        {
          "id": 961592,
          "postDate": "2020-08-07T10:16:53.430Z",
          "content": "<p>Gilles, I do not agree. At least not fully. Here is why.</p>\n<p>It is often the case that the most important are actually those negative that turn into false positive. You simply do not wanna falsely indicate presence of a condition such as a disease especially when a treatment also causes harm. Even if not harmful directly, it's still super stressful and should be avoided. <em>Primum non nocere</em></p>",
          "rawMarkdown": "Gilles, I do not agree. At least not fully. Here is why.\n\nIt is often the case that the most important are actually those negative that turn into false positive. You simply do not wanna falsely indicate presence of a condition such as a disease especially when a treatment also causes harm. Even if not harmful directly, it's still super stressful and should be avoided. *Primum non nocere*",
          "votes": 1
        },
        {
          "id": 961603,
          "postDate": "2020-08-07T10:24:36.413Z",
          "content": "<p>It's mostly much worse the other way around though. Diagnosing a healthy person as sick will cost an unnecessary operation and possibly some harm. Diagnosing a sick person as healthy (e.g. cancer) will cost a human life.</p>",
          "rawMarkdown": "It's mostly much worse the other way around though. Diagnosing a healthy person as sick will cost an unnecessary operation and possibly some harm. Diagnosing a sick person as healthy (e.g. cancer) will cost a human life.",
          "votes": 3
        },
        {
          "id": 967324,
          "postDate": "2020-08-12T06:57:17.480Z",
          "content": "<p>It means we want to focus more on false-negative rather than false-positive.</p>",
          "rawMarkdown": "It means we want to focus more on false-negative rather than false-positive."
        },
        {
          "id": 967329,
          "postDate": "2020-08-12T07:02:33.480Z",
          "content": "<p>Ya <a href=\"https://www.kaggle.com/fangao\" target=\"_blank\">@fangao</a> but how then our model will learn minority class from such a small number of samples?<br>\nAre the samples for minority class sufficient for learning?</p>",
          "rawMarkdown": "Ya @fangao but how then our model will learn minority class from such a small number of samples?\nAre the samples for minority class sufficient for learning?"
        }
      ]
    },
    {
      "id": 961919,
      "postDate": "2020-08-07T15:54:34.447Z",
      "content": "<h3>Confusion between \"upsample\" and \"add data\"</h3>\n<blockquote>\n  <p>There is some confusion about wording however. Upsampling the minority class means that samples form minority class are sampled more than once (duplicated). Obviously. some confuse that with adding new samples for the minority class.</p>\n</blockquote>\n<p>When using TFRecords, many participants \"upsample\" by \"adding more TFRecords\" that contain only malignant images that are already in their train data. This is causing the confusion between the words \"upsample\" and \"add data\". I explain how to \"upsample\" by \"adding data\" <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a> and share TFRecords containing only malignant. </p>\n<blockquote>\n  <p>Adding more data helps in general. Adding only samples from minority class does not help necessarily however. </p>\n</blockquote>\n<p>If we add samples from the minority class that are already present in our train, then we are \"upsampling\". If the samples from the minority class are not already present in our train, we are \"not upsampling\". If we wish to experiment with \"upsampling\" 2019 external data, then we first add my 2019 external dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a> and second add my 2019 malignant dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>. (To \"upsample\" 2020 data, we add my 2020 malignant <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>).</p>\n<h3>Example \"upsample\" Notebook</h3>\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\" target=\"_blank\">here</a> that adds 2019 data and then \"upsamples\" the 2019 data. It also \"upsamples\" the 2020 data. Both \"upsample\" by \"adding TFRecords\". The notebook uses 128x128 images and EfficientNetB0. It scores 0.910 \"CV\" and 0.916 LB. Note how well the CV and LB align. The \"CV\" score is computed by ensembling the 3 experiments as follows:</p>\n<pre><code>from sklearn.metric import roc_auc_score\noof = df_oof.iloc[:6552,2].values + df_oof.iloc[6552:6552*2,2].values\\\n          + df_oof.iloc[6552*2:,2].values\nprint('Overall CV AUC =', roc_auc_score(df.iloc[:6552,1].values,oof) )\n# Overall CV AUC = 0.909766721673346\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F557caa6c4821d8ed7f3644cadd43766c%2Flb.png?generation=1596815325323718&amp;alt=media\" alt=\"\"> </p>",
      "rawMarkdown": "### Confusion between \"upsample\" and \"add data\"\n&gt; There is some confusion about wording however. Upsampling the minority class means that samples form minority class are sampled more than once (duplicated). Obviously. some confuse that with adding new samples for the minority class.\n\nWhen using TFRecords, many participants \"upsample\" by \"adding more TFRecords\" that contain only malignant images that are already in their train data. This is causing the confusion between the words \"upsample\" and \"add data\". I explain how to \"upsample\" by \"adding data\" [here][1] and share TFRecords containing only malignant. \n\n&gt; Adding more data helps in general. Adding only samples from minority class does not help necessarily however. \n\nIf we add samples from the minority class that are already present in our train, then we are \"upsampling\". If the samples from the minority class are not already present in our train, we are \"not upsampling\". If we wish to experiment with \"upsampling\" 2019 external data, then we first add my 2019 external dataset [here][2] and second add my 2019 malignant dataset [here][1]. (To \"upsample\" 2020 data, we add my 2020 malignant [here][1]).\n\n### Example \"upsample\" Notebook\n\nI posted a starter notebook [here][3] that adds 2019 data and then \"upsamples\" the 2019 data. It also \"upsamples\" the 2020 data. Both \"upsample\" by \"adding TFRecords\". The notebook uses 128x128 images and EfficientNetB0. It scores 0.910 \"CV\" and 0.916 LB. Note how well the CV and LB align. The \"CV\" score is computed by ensembling the 3 experiments as follows:\n\n    from sklearn.metric import roc_auc_score\n    oof = df_oof.iloc[:6552,2].values + df_oof.iloc[6552:6552*2,2].values\\\n              + df_oof.iloc[6552*2:,2].values\n    print('Overall CV AUC =', roc_auc_score(df.iloc[:6552,1].values,oof) )\n    # Overall CV AUC = 0.909766721673346\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F557caa6c4821d8ed7f3644cadd43766c%2Flb.png?generation=1596815325323718&amp;alt=media) \n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[3]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
      "votes": 3,
      "replies": [
        {
          "id": 961939,
          "postDate": "2020-08-07T16:25:44.127Z",
          "content": "<blockquote>\n  <p>Why Confusion Exists</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Indeed, it probably comes from your posts where you both upsample and add new data.  </p>\n<p>As I wrote, terminology isn't that important as long as we speak about the same.  Given the first comments above where about adding data rather than upsampling, I thought I had to make it clearer.  I hope we are all set now.</p>",
          "rawMarkdown": "&gt; Why Confusion Exists\n\n@cdeotte Indeed, it probably comes from your posts where you both upsample and add new data.  \n\nAs I wrote, terminology isn't that important as long as we speak about the same.  Given the first comments above where about adding data rather than upsampling, I thought I had to make it clearer.  I hope we are all set now.",
          "replies": [
            {
              "id": 961940,
              "postDate": "2020-08-07T16:26:29.910Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 964238,
      "postDate": "2020-08-09T17:15:31.043Z",
      "content": "<p>I edited the post, given upsampling seems to enable faster convergence.  This alone justifies upsampling IMHO.  Thanks to all who shared this.</p>",
      "rawMarkdown": "I edited the post, given upsampling seems to enable faster convergence.  This alone justifies upsampling IMHO.  Thanks to all who shared this.",
      "votes": 1,
      "replies": [
        {
          "id": 964359,
          "postDate": "2020-08-09T19:18:09.403Z",
          "content": "<p>It must require some retuning because I cannot match the roc-auc I got without upsampling.</p>",
          "rawMarkdown": "It must require some retuning because I cannot match the roc-auc I got without upsampling."
        }
      ]
    },
    {
      "id": 963731,
      "postDate": "2020-08-09T08:37:49.087Z",
      "content": "<p>In my brief experience, with a more balanced ratio between the classes you can converge in fewer steps</p>",
      "rawMarkdown": "In my brief experience, with a more balanced ratio between the classes you can converge in fewer steps",
      "votes": 1,
      "replies": [
        {
          "id": 964006,
          "postDate": "2020-08-09T14:07:54.663Z",
          "content": "<p>indeed, several other people made the same point.  Thanks for adding more confirmation.  I will certainly try at a  point.</p>\n<p>Does it mean you decreased the number of epochs ?</p>",
          "rawMarkdown": "indeed, several other people made the same point.  Thanks for adding more confirmation.  I will certainly try at a  point.\n\nDoes it mean you decreased the number of epochs ?"
        }
      ]
    },
    {
      "id": 963427,
      "postDate": "2020-08-09T03:07:58.640Z",
      "content": "<p>Just a thought... if upsampling without Augmentation, a commonsense tells me it's doing no more than just creating duplicates. However, if upsampling with extensive augmentation, wouldn't the problem be a little bit more debatable? As augmentation is adding variance, which is not necessarily a bad thing to have. This also explains why it makes model converges faster as model is seeing more variations of one particular malignant picture at one epoch. </p>",
      "rawMarkdown": "Just a thought... if upsampling without Augmentation, a commonsense tells me it's doing no more than just creating duplicates. However, if upsampling with extensive augmentation, wouldn't the problem be a little bit more debatable? As augmentation is adding variance, which is not necessarily a bad thing to have. This also explains why it makes model converges faster as model is seeing more variations of one particular malignant picture at one epoch. ",
      "votes": 1,
      "replies": [
        {
          "id": 964007,
          "postDate": "2020-08-09T14:08:50.757Z",
          "content": "<p>This point was made already as well.  Question still remains: do you get better roc-auc that way?</p>",
          "rawMarkdown": "This point was made already as well.  Question still remains: do you get better roc-auc that way?"
        }
      ]
    },
    {
      "id": 962977,
      "postDate": "2020-08-08T15:22:57.543Z",
      "content": "<p>Thank you for discussing openly, freely and without tickets. Development and the next best level take place through these analyzes. I listen and take notes.</p>",
      "rawMarkdown": "Thank you for discussing openly, freely and without tickets. Development and the next best level take place through these analyzes. I listen and take notes.",
      "votes": 1,
      "replies": [
        {
          "id": 963211,
          "postDate": "2020-08-08T18:54:23.613Z",
          "content": "<p>Thank you.  Glad to see that some don't mind me asking questions around ;)</p>",
          "rawMarkdown": "Thank you.  Glad to see that some don't mind me asking questions around ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 961667,
      "postDate": "2020-08-07T11:29:30.040Z",
      "content": "<p>I updated the topic to address the main source of disagreement: here upsampling has its usual meaning, where samples from the minority class are replicated.</p>",
      "rawMarkdown": "I updated the topic to address the main source of disagreement: here upsampling has its usual meaning, where samples from the minority class are replicated.",
      "votes": 1
    },
    {
      "id": 961539,
      "postDate": "2020-08-07T09:15:44.373Z",
      "content": "<p>Can it help with extreme or specific upsample-TTA to the upsampled data, does it help or status quo.<br>\nWhen upsample some classes, can it increase training/class weight for some samples and maybe cause a new overfitting when trying to avoid it in the first place? <br>\nSolution(?): Maybe one should use both upsamples and downsample together, do little of both, midway solution and with different amount of +/- samples each process in the same CV fold or the next, same data but different class weight every time.</p>",
      "rawMarkdown": "Can it help with extreme or specific upsample-TTA to the upsampled data, does it help or status quo.\nWhen upsample some classes, can it increase training/class weight for some samples and maybe cause a new overfitting when trying to avoid it in the first place? \nSolution(?): Maybe one should use both upsamples and downsample together, do little of both, midway solution and with different amount of +/- samples each process in the same CV fold or the next, same data but different class weight every time.",
      "votes": 1
    },
    {
      "id": 961445,
      "postDate": "2020-08-07T07:33:40.467Z",
      "content": "<p>The only reason I am concerned about target balancing here is deep learning and batches (contrary to boosting models that have a broad view of global statistics).</p>\n<p>We have 2% of positive samples, so if because of model size and image size I can’t use a bigger batch size than 8 or 16, I’ll end up training my model with a lot of batches only containing 0 targets and a few of them containing 1 or 2 ones.<br>\nIt feels like this could hurt the gradient descend or at least slow it down.</p>\n<p>That’s why trying to have at least one positive example per batch could seem fair, or giving more importance to ones so that they end up having a word to say when they show up in one batch.</p>\n<p>Here I agree that it might not drastically change the final score, but if you can get the same result with 10 epochs instead of 20 by doing something about class imbalance then you save a lot of time for more experiments.</p>",
      "rawMarkdown": "The only reason I am concerned about target balancing here is deep learning and batches (contrary to boosting models that have a broad view of global statistics).\n\nWe have 2% of positive samples, so if because of model size and image size I can’t use a bigger batch size than 8 or 16, I’ll end up training my model with a lot of batches only containing 0 targets and a few of them containing 1 or 2 ones.\nIt feels like this could hurt the gradient descend or at least slow it down.\n\nThat’s why trying to have at least one positive example per batch could seem fair, or giving more importance to ones so that they end up having a word to say when they show up in one batch.\n\nHere I agree that it might not drastically change the final score, but if you can get the same result with 10 epochs instead of 20 by doing something about class imbalance then you save a lot of time for more experiments.",
      "votes": 1,
      "replies": [
        {
          "id": 961580,
          "postDate": "2020-08-07T10:00:15.297Z",
          "content": "<p>Perhaps you can mitigate this with gradient accumulation?</p>",
          "rawMarkdown": "Perhaps you can mitigate this with gradient accumulation?",
          "votes": 5
        },
        {
          "id": 961955,
          "postDate": "2020-08-07T16:49:22.477Z",
          "content": "<p>yes that's one solution as well! I've never used it but definitely something to try!</p>",
          "rawMarkdown": "yes that's one solution as well! I've never used it but definitely something to try!",
          "votes": 1
        }
      ]
    },
    {
      "id": 961105,
      "postDate": "2020-08-06T23:06:36.070Z",
      "content": "<p>How are you to predict \"malignant\" when you don't have many malignant to train from?  For the order to be correct, you must be able to determine benign vs malignant.  You can't possibly determine malignant unless you have enough data to help you do that.</p>",
      "rawMarkdown": "How are you to predict \"malignant\" when you don't have many malignant to train from?  For the order to be correct, you must be able to determine benign vs malignant.  You can't possibly determine malignant unless you have enough data to help you do that.",
      "votes": 1,
      "replies": [
        {
          "id": 961120,
          "postDate": "2020-08-06T23:53:35.833Z",
          "content": "<p>You mix using more data, eg external data, and fixing class imbalance I'm afraid.</p>",
          "rawMarkdown": "You mix using more data, eg external data, and fixing class imbalance I'm afraid.",
          "votes": 2
        },
        {
          "id": 961131,
          "postDate": "2020-08-07T00:03:13.080Z",
          "content": "<p>To be clearer maybe: upsampling from the same data does not add new data.  Same for class weights.  I didn't see improvement when doing it, and I am asking if other see improvement.</p>",
          "rawMarkdown": "To be clearer maybe: upsampling from the same data does not add new data.  Same for class weights.  I didn't see improvement when doing it, and I am asking if other see improvement.",
          "votes": 2
        },
        {
          "id": 961205,
          "postDate": "2020-08-07T02:08:10.307Z",
          "content": "<p>What if you use a heavy set of augmentations? Then upsampling “sort of” adds new data, since (almost) all extra images are going to be unique. </p>\n\n<p>In the end, upsampling increases the share of malignant images in every mini batch. In my understanding, this can help to learn patters from malignant images better, since their contribution to the gradients you compute is higher compared to an original batch with a fewer malignant. </p>\n\n<p>That said, my results so far are  mixed.</p>",
          "rawMarkdown": "What if you use a heavy set of augmentations? Then upsampling “sort of” adds new data, since (almost) all extra images are going to be unique. \n\nIn the end, upsampling increases the share of malignant images in every mini batch. In my understanding, this can help to learn patters from malignant images better, since their contribution to the gradients you compute is higher compared to an original batch with a fewer malignant. \n\nThat said, my results so far are  mixed.",
          "votes": 2,
          "replies": [
            {
              "id": 961657,
              "postDate": "2020-08-07T11:16:09.363Z",
              "content": "<blockquote>\n  <p>What if you use a heavy set of augmentations? Then upsampling “sort of” adds new data, since (almost) all extra images are going to be unique. </p>\n</blockquote>\n<p>If you use the same set of augmentations for positive and negative then you don't change class proportion.  This is not upsampling the minority class.</p>",
              "rawMarkdown": "&gt; What if you use a heavy set of augmentations? Then upsampling “sort of” adds new data, since (almost) all extra images are going to be unique. \n\nIf you use the same set of augmentations for positive and negative then you don't change class proportion.  This is not upsampling the minority class.",
              "votes": 1
            }
          ]
        },
        {
          "id": 961353,
          "postDate": "2020-08-07T05:42:44.753Z",
          "content": "<p>I have not seen any improvement by using Upsampling, by upsampling I mean showing the malignant images more number of times in each epoch then they are. But I was expecting it to work since there are few examples in each batch containing malignant the model may not get enough chance to know about it. But from some of my experimentation, it is not helping. I have a few more ideas that I wanted to try. Hope I can make it work. </p>",
          "rawMarkdown": "I have not seen any improvement by using Upsampling, by upsampling I mean showing the malignant images more number of times in each epoch then they are. But I was expecting it to work since there are few examples in each batch containing malignant the model may not get enough chance to know about it. But from some of my experimentation, it is not helping. I have a few more ideas that I wanted to try. Hope I can make it work. ",
          "votes": 1
        },
        {
          "id": 967340,
          "postDate": "2020-08-12T07:09:40.323Z",
          "content": "<p>Maybe it is overfitting on the malignant images. And the model is not able to generalise from the same small set of malignant images. </p>",
          "rawMarkdown": "Maybe it is overfitting on the malignant images. And the model is not able to generalise from the same small set of malignant images. "
        }
      ]
    },
    {
      "id": 962796,
      "postDate": "2020-08-08T12:53:00.953Z",
      "content": "<p>For what it's worth, this is what I'm currently trying with neural nets, although I can't claim any particular success with it. My nets train, loss drops, and predictions are much better than random chance but very far from medal territory:</p>\n\n<ol>\n<li><p>I'm working on an x86_64 workstation running Linux, R, Keras, Tensorflow, CUDA, and an Nvidia Quadro GP100 GPU.</p></li>\n<li><p>I'm using no external data, just the original JPEG images resized to 384x384. But I also populate some additional input channels besides the 3 RGB with a few other features computed from the images and meta data, each such feature replicating its value across every pixel of its associated channel.</p></li>\n<li><p>I'm using a home-grown neural net of no particular distinction, containing 4 convolutional layers and associated parametric rectified linear activation, batch normalization, and max pooling layers. I flatten the output of those, feed the result to a dense layer with a single output value, and apply a sigmoid activation. I use BCE with some label smoothing as my loss function, and an rmsprop optimizer.</p></li>\n<li><p>I don't train in \"epochs\", but in a long series of individual batches using Keras's \"train_on_batch\" function. Each batch contains 24 samples chosen randomly from the training data, but 3 slots (12.5%) are reserved for positive examples and the remaining 21 for negative. This amounts to a considerable boost in the percentage of positive examples compared to the 1.76% in the training data. To each image in the batch I randomly apply, with specified probabilities, some combination of augmentations chosen currently from \"flip\", \"flop\", and \"transpose\".</p></li>\n<li><p>As training continues it eventually gets stuck making no additional progress reducing the loss. When this happens my algorithm divides the learning rate in half, restores the net's weights from before progress halted, and continues training. After several such restarts (currently 6), if progress halts again training is terminated. This typically happens between 10000 and 15000 batches, which can take several hours to run.</p></li>\n<li><p>I use TTA (test time augmentation) by predicting for all eight combinations of original, flipped, flopped, and/or transposed images, and averaging the results together.</p></li>\n<li><p>To try to reduce overfitting I conduct several runs with different random number seeds and ensemble the results by averaging their predictions after converting them to ranks.</p></li>\n</ol>",
      "rawMarkdown": "For what it's worth, this is what I'm currently trying with neural nets, although I can't claim any particular success with it. My nets train, loss drops, and predictions are much better than random chance but very far from medal territory:\n\n1. I'm working on an x86_64 workstation running Linux, R, Keras, Tensorflow, CUDA, and an Nvidia Quadro GP100 GPU.\n\n2. I'm using no external data, just the original JPEG images resized to 384x384. But I also populate some additional input channels besides the 3 RGB with a few other features computed from the images and meta data, each such feature replicating its value across every pixel of its associated channel.\n\n3. I'm using a home-grown neural net of no particular distinction, containing 4 convolutional layers and associated parametric rectified linear activation, batch normalization, and max pooling layers. I flatten the output of those, feed the result to a dense layer with a single output value, and apply a sigmoid activation. I use BCE with some label smoothing as my loss function, and an rmsprop optimizer.\n\n4. I don't train in \"epochs\", but in a long series of individual batches using Keras's \"train_on_batch\" function. Each batch contains 24 samples chosen randomly from the training data, but 3 slots (12.5%) are reserved for positive examples and the remaining 21 for negative. This amounts to a considerable boost in the percentage of positive examples compared to the 1.76% in the training data. To each image in the batch I randomly apply, with specified probabilities, some combination of augmentations chosen currently from \"flip\", \"flop\", and \"transpose\".\n\n5. As training continues it eventually gets stuck making no additional progress reducing the loss. When this happens my algorithm divides the learning rate in half, restores the net's weights from before progress halted, and continues training. After several such restarts (currently 6), if progress halts again training is terminated. This typically happens between 10000 and 15000 batches, which can take several hours to run.\n\n6. I use TTA (test time augmentation) by predicting for all eight combinations of original, flipped, flopped, and/or transposed images, and averaging the results together.\n\n7. To try to reduce overfitting I conduct several runs with different random number seeds and ensemble the results by averaging their predictions after converting them to ranks.",
      "votes": 2,
      "replies": [
        {
          "id": 963212,
          "postDate": "2020-08-08T18:55:32.087Z",
          "content": "<p>Have you compared with a lower percentage of positives?</p>",
          "rawMarkdown": "Have you compared with a lower percentage of positives?"
        },
        {
          "id": 963277,
          "postDate": "2020-08-08T20:43:20.353Z",
          "content": "<p>Not yet with my current setup, but I may try that, as well as larger %.  I've also been working on other approaches to model-building for this contest, such as lightgbm with a bunch of features computed in R, as well as a rather exotic genetic algorithm specialized for image classification (perhaps to be described in a later post).  So my computers and my own brain have been pretty busy lately.</p>",
          "rawMarkdown": "Not yet with my current setup, but I may try that, as well as larger %.  I've also been working on other approaches to model-building for this contest, such as lightgbm with a bunch of features computed in R, as well as a rather exotic genetic algorithm specialized for image classification (perhaps to be described in a later post).  So my computers and my own brain have been pretty busy lately.",
          "votes": 1
        }
      ]
    },
    {
      "id": 962784,
      "postDate": "2020-08-08T12:44:21.673Z",
      "content": "<p>While working on a classification problem I noticed that data is slightly imbalanced. So I used data balancing techniques like UpSample and DownSample. \nIn UpSampling data points of minority count is are increased to the level of data points of majority count. It simply creates duplicate data points in the range of same data points. \nIn DowSampling inverse process is done. Some data points from class of majority count are removed in order to match the count of class of minority count.</p>\n\n<p>Find complete work here\n<a href=\"https://www.kaggle.com/omkargurav/loan-approval-prediction\">Loan Approval Prediction</a></p>\n\n<p>`A = list(data.Loan_Status).count(1)\nB = list(data.Loan_Status).count(0)\nprint(\"Count of 1: \",A,\"\\nCount of 0: \",B)</p>\n\n<p>fig = px.bar((A,B),x=[\"Approved\",\"Rejected\"],y=[A,B],color=[A,B])\nfig.show()\nCount of 1:  422 \nCount of 0:  192`</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3559882%2F101b8f6e6a6df24c5ed31afea98318bb%2Fnewplot.png?generation=1596892428969836&amp;alt=media\" alt=\"\"></p>\n\n<p>`#Getting seperated data with 1 and 0 status.\ndf_majority = new_data[new_data.Loan_Status==1]\ndf_minority = new_data[new_data.Loan_Status==0]</p>\n\n<h1>Here we are downsampling the Majority Class Data Points.</h1>\n\n<h1>i.e. We will get equal amount of datapoint as Minority class from Majority class</h1>\n\n<p>df_manjority_downsampled = resample(df_majority,replace=False,n_samples=192,random_state=123)\ndf_downsampled = pd.concat([df_manjority_downsampled,df_minority])\nprint(\"Downsampled data:-&gt;\\n\",df_downsampled.Loan_Status.value_counts())</p>\n\n<h1>Here we are upsampling the Minority Class Data Points.</h1>\n\n<h1>i.e. We will get equal amount of datapoint as Majority class from Minority class</h1>\n\n<p>df_monority_upsampled = resample(df_minority,replace=True,n_samples=422,random_state=123)\ndf_upsampled = pd.concat([df_majority,df_monority_upsampled])\nprint(\"Upsampled data:-&gt;\\n\",df_upsampled.Loan_Status.value_counts())\nDownsampled data:-&gt;\n 1    192\n0    192\nName: Loan_Status, dtype: int64\nUpsampled data:-&gt;\n 1    422\n0    422\nName: Loan_Status, dtype: int64`</p>\n\n<p>After this method accuracy increased by around 8-10%.</p>",
      "rawMarkdown": "While working on a classification problem I noticed that data is slightly imbalanced. So I used data balancing techniques like UpSample and DownSample. \nIn UpSampling data points of minority count is are increased to the level of data points of majority count. It simply creates duplicate data points in the range of same data points. \nIn DowSampling inverse process is done. Some data points from class of majority count are removed in order to match the count of class of minority count.\n\nFind complete work here\n[Loan Approval Prediction](https://www.kaggle.com/omkargurav/loan-approval-prediction)\n\n`A = list(data.Loan_Status).count(1)\nB = list(data.Loan_Status).count(0)\nprint(\"Count of 1"
    },
    {
      "id": 962240,
      "postDate": "2020-08-08T00:21:55.820Z",
      "content": "<p>In my case, balancing classes prevented me from getting results with the interference of one majority of data type.</p>",
      "rawMarkdown": "In my case, balancing classes prevented me from getting results with the interference of one majority of data type."
    },
    {
      "id": 961727,
      "postDate": "2020-08-07T13:03:02.980Z",
      "content": "<p>&gt; I have not found it to be useful here, and my experience is that it is never useful when using deep learning</p>\n\n<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>It does help in some case, like in <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/overview\">Flower Classification with TPUs</a></p>\n\n<p>As shown in my kernel <a href=\"https://www.kaggle.com/yihdarshieh/detailed-guide-to-custom-training-with-tpus#no-oversampling-vs.-oversampling-%24N=300%24\">Detailed guide to custom training with TPUs</a>, oversampling improves on the validation dataset.</p>\n\n<pre><code>without oversampling\n    valid recall: 0.890801499933458\n    valid precision: 0.9074883670834412\n    valid f1: 0.8926506729634966\n\nwith oversampling\n    valid recall: 0.9269954377916817\n    valid precision: 0.9235320148160663\n    valid f1: 0.9227047892043186\n</code></pre>\n\n<p>Even training with oversampling but fewer epochs (so equivalent number of examples trained) is slightly better</p>\n\n<pre><code>    valid recall: 0.9156978000819738\n    valid precision: 0.8914945977622719\n    valid f1: 0.8997503256064424\n</code></pre>\n\n<p>I didn't submit the results to check on the test dataset for the LB score though.</p>",
      "rawMarkdown": "&gt; I have not found it to be useful here, and my experience is that it is never useful when using deep learning\n\n@cpmpml \n\nIt does help in some case, like in [Flower Classification with TPUs](https://www.kaggle.com/c/flower-classification-with-tpus/overview)\n\nAs shown in my kernel [Detailed guide to custom training with TPUs](https://www.kaggle.com/yihdarshieh/detailed-guide-to-custom-training-with-tpus#no-oversampling-vs.-oversampling-$N=300$), oversampling improves on the validation dataset.\n\n    without oversampling\n        valid recall: 0.890801499933458\n        valid precision: 0.9074883670834412\n        valid f1: 0.8926506729634966\n\n    with oversampling\n        valid recall: 0.9269954377916817\n        valid precision: 0.9235320148160663\n        valid f1: 0.9227047892043186\n\nEven training with oversampling but fewer epochs (so equivalent number of examples trained) is slightly better\n\n        valid recall: 0.9156978000819738\n        valid precision: 0.8914945977622719\n        valid f1: 0.8997503256064424\n\nI didn't submit the results to check on the test dataset for the LB score though.",
      "replies": [
        {
          "id": 961736,
          "postDate": "2020-08-07T13:13:58.153Z",
          "content": "<p>In Flower competition, metric is F1 score, which depends on prediction values.  Here the metric is roc-auc, which is independent of prediction values.  roc-auc only depends on the ordering.</p>\n<p>I am not surprised upsampling helps with F1 score, as it leads to better probability calibration.  I hinted at in my topic BTW.  But I have yet to see how upsampling helps for roc-auc.</p>",
          "rawMarkdown": "In Flower competition, metric is F1 score, which depends on prediction values.  Here the metric is roc-auc, which is independent of prediction values.  roc-auc only depends on the ordering.\n\nI am not surprised upsampling helps with F1 score, as it leads to better probability calibration.  I hinted at in my topic BTW.  But I have yet to see how upsampling helps for roc-auc.\n\n",
          "votes": 4
        },
        {
          "id": 961739,
          "postDate": "2020-08-07T13:16:01.240Z",
          "content": "<p>If you would tune F1 thresholds for cutoff, you would get similar boosts. Upsampling - as <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> says - moves the probabilities, so default cutoff might be better.</p>",
          "rawMarkdown": "If you would tune F1 thresholds for cutoff, you would get similar boosts. Upsampling - as @cpmpml says - moves the probabilities, so default cutoff might be better.",
          "votes": 3
        },
        {
          "id": 961744,
          "postDate": "2020-08-07T13:19:42.160Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I didn't realize that you are talking for roc-auc rather than in general. Sorry to bother.</p>",
          "rawMarkdown": "@cpmpml I didn't realize that you are talking for roc-auc rather than in general. Sorry to bother.",
          "votes": 4,
          "replies": [
            {
              "id": 961754,
              "postDate": "2020-08-07T13:30:14.660Z",
              "content": "<p>No need to be sorry, you shared information ;)</p>",
              "rawMarkdown": "No need to be sorry, you shared information ;)\n",
              "votes": 2
            }
          ]
        },
        {
          "id": 961924,
          "postDate": "2020-08-07T16:07:51.807Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> , do you think that upsampling could only calibrate probabilities but not able to make the model have more corrected predictions?</p>",
          "rawMarkdown": "@cpmpml , do you think that upsampling could only calibrate probabilities but not able to make the model have more corrected predictions?"
        },
        {
          "id": 961950,
          "postDate": "2020-08-07T16:46:05.983Z",
          "content": "<p>In my experience, it's mostly a trade-off between class-specific performances. Most of the times, you sacrifice a bit of performance in the majority class for more performance in the minority class (which is often more relevant). Nevertheless, there are use cases where the global accuracy also increased by oversampling. Moreover, I could (probably) draw artificial examples where, for example. SMOTE would be beneficial for both classes. This is because the biases/assumptions introduced by the over-sampling methods actually helped drawing better decision boundaries. These are, however, very rare cases.</p>",
          "rawMarkdown": "In my experience, it's mostly a trade-off between class-specific performances. Most of the times, you sacrifice a bit of performance in the majority class for more performance in the minority class (which is often more relevant). Nevertheless, there are use cases where the global accuracy also increased by oversampling. Moreover, I could (probably) draw artificial examples where, for example. SMOTE would be beneficial for both classes. This is because the biases/assumptions introduced by the over-sampling methods actually helped drawing better decision boundaries. These are, however, very rare cases.",
          "replies": [
            {
              "id": 961970,
              "postDate": "2020-08-07T17:04:50.210Z",
              "content": "<blockquote>\n  <p>it's mostly a trade-off between class-specific performances.</p>\n</blockquote>\n<p>What does it mean when roc-auc is the metric?  </p>\n<p>Again, I know upsampling helps when the metric depends on the absolute value of predictions.  We are not in this case here.</p>",
              "rawMarkdown": "&gt; it's mostly a trade-off between class-specific performances.\n\nWhat does it mean when roc-auc is the metric?  \n\nAgain, I know upsampling helps when the metric depends on the absolute value of predictions.  We are not in this case here."
            }
          ]
        }
      ]
    },
    {
      "id": 961194,
      "postDate": "2020-08-07T01:44:47.600Z",
      "content": "<p>It may not necessarily be related to this thread, but I think the necessary goal for this kind of applications is to be able to correctly classify all malignant cases without erring too much on benign ones. I don't know if the metric used is the best for that ... even if it is for the competition ... will it be the best for the use of the knowledge obtained?</p>",
      "rawMarkdown": "It may not necessarily be related to this thread, but I think the necessary goal for this kind of applications is to be able to correctly classify all malignant cases without erring too much on benign ones. I don't know if the metric used is the best for that ... even if it is for the competition ... will it be the best for the use of the knowledge obtained?"
    },
    {
      "id": 961130,
      "postDate": "2020-08-07T00:02:02.747Z",
      "content": "<p>Class imbalance is a real issue and we should at least try to remedy it. In my experience, class weights work sometimes but aren't reliable. You said you also tried upsampling the positive classes and it didn't work, but it really depends on what specific upsampling methods you used. If you are not convinced, try training on external melanoma images (effectively upsampling the positive class) and see if your scores improve.</p>",
      "rawMarkdown": "Class imbalance is a real issue and we should at least try to remedy it. In my experience, class weights work sometimes but aren't reliable. You said you also tried upsampling the positive classes and it didn't work, but it really depends on what specific upsampling methods you used. If you are not convinced, try training on external melanoma images (effectively upsampling the positive class) and see if your scores improve.",
      "replies": [
        {
          "id": 961133,
          "postDate": "2020-08-07T00:03:43.477Z",
          "content": "<blockquote>\n  <p>Class imbalance is a real issue</p>\n</blockquote>\n<p>Why is it a real issue?</p>",
          "rawMarkdown": "&gt; Class imbalance is a real issue\n\nWhy is it a real issue?",
          "votes": 2
        },
        {
          "id": 961134,
          "postDate": "2020-08-07T00:04:41.783Z",
          "content": "<blockquote>\n  <p>try training on external melanoma images (effectively upsampling the positive class) and see if your scores improve.</p>\n</blockquote>\n<p>That's not upsampling.  That's using more data.</p>",
          "rawMarkdown": "&gt;  try training on external melanoma images (effectively upsampling the positive class) and see if your scores improve.\n\nThat's not upsampling.  That's using more data.\n\n",
          "votes": 2
        },
        {
          "id": 961140,
          "postDate": "2020-08-07T00:09:08.217Z",
          "content": "<p>Yes, sorry I misspoke but it still helps address class imbalance, and there are notebooks here showing that adding external data improved their scores (like Chris Deotte).</p>",
          "rawMarkdown": "Yes, sorry I misspoke but it still helps address class imbalance, and there are notebooks here showing that adding external data improved their scores (like Chris Deotte).",
          "replies": [
            {
              "id": 961141,
              "postDate": "2020-08-07T00:11:13.280Z",
              "content": "<p>Again, using more data helps, no argument here.  I am discussing upsampling, which means that some samples are seen more often than others when training.</p>",
              "rawMarkdown": "Again, using more data helps, no argument here.  I am discussing upsampling, which means that some samples are seen more often than others when training.",
              "votes": 1
            }
          ]
        },
        {
          "id": 961153,
          "postDate": "2020-08-07T00:28:49.713Z",
          "content": "<p>Just a point.  Literally just about anyone upsampling is also doing augmentation (<em>corrected, I had put TTA</em>).  Which means each upsampled image is unique.......possibly very unique: cutout, saturation, contrast, brightness, rotation, skew, etc, etc, etc.  Whic means, when we talk of upsampling, we are talking about new synthetic data.</p>",
          "rawMarkdown": "Just a point.  Literally just about anyone upsampling is also doing augmentation (*corrected, I had put TTA*).  Which means each upsampled image is unique.......possibly very unique: cutout, saturation, contrast, brightness, rotation, skew, etc, etc, etc.  Whic means, when we talk of upsampling, we are talking about new synthetic data.",
          "votes": 4,
          "replies": [
            {
              "id": 961183,
              "postDate": "2020-08-07T01:26:12.253Z",
              "content": "<blockquote>\n  <p>Just a point. Literally just about anyone upsampling is also doing TTA. Which means each upsampled image is unique…….possibly very unique: cutout, saturation, contrast, brightness, rotation, skew, etc, etc, etc. Which means, when we talk of upsampling, we are talking about new synthetic data.</p>\n</blockquote>\n<p>TTA is test time augmentation.  At this stage you don't use target values, hence you perform the same TTA for negative and positive samples.  You do not change class imbalance via TTA.</p>\n<p>I wish someone was answering my question instead of discussing other topics.</p>",
              "rawMarkdown": "&gt; Just a point. Literally just about anyone upsampling is also doing TTA. Which means each upsampled image is unique…….possibly very unique: cutout, saturation, contrast, brightness, rotation, skew, etc, etc, etc. Which means, when we talk of upsampling, we are talking about new synthetic data.\n\nTTA is test time augmentation.  At this stage you don't use target values, hence you perform the same TTA for negative and positive samples.  You do not change class imbalance via TTA.\n\nI wish someone was answering my question instead of discussing other topics.\n\n",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 967068,
      "postDate": "2020-08-11T22:15:47.450Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 961492,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-07T08:26:22.490000",
      "content": "<p>Doing nothing in imbalanced problems is often the best.</p>",
      "votes": 18,
      "replies": []
    },
    {
      "id": 962051,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-08-07T18:43:56.867000",
      "content": "<p>Small Example I want to add. Maybe related indirectly. Noisy student paper did quite a good experiment on if it is essential to balance data or not when adding external data. In short, they :<br>\n1) Trained Effnet On Imagenet<br>\n2) Predicted Random Images from Interner(out of distribution from original Imagenet Dataset) -&gt; generate Pseudolabels<br>\n3) Added this Images and retrained further. </p>\n<p>Now when you predict Images that are not present in the imagenet dataset, you will get different class distribution. The question that was raised is essential to balance classes when retraining network? </p>\n<p>Experimental Answer: For small model -Yes! (you will gain slight boost). It was not helpful for bigger models since they have more power to learn and separate features. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2F9a949952781f1d7da2412c95cdc8faf2%2FScreen%20Shot%202020-08-07%20at%202.41.15%20PM.png?generation=1596825750152331&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2Ff96afd76d6c8301316bdae17d6f020a2%2FScreen%20Shot%202020-08-07%20at%202.41.20%20PM.png?generation=1596825764119842&amp;alt=media\" alt=\"\"></p>\n<p>As I mentioned its probably very remotely related to the topic but perhaps can give a good experimental evidence for the discussion. </p>\n<p>Paper (<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a>)</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 961490,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-08-07T08:24:03.537000",
      "content": "<p>I don't upsample but I do some  balancing in some of my experiments for the sole purpose to speed up training. <br>\nFor that, I discard some benign ones , but it's done online i.e when the batch is being fed to the model. Because model don't need to focus that much on benign cases and this save training time too ;)</p>",
      "votes": 9,
      "replies": [
        {
          "id": 961531,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-08-07T09:05:33.227000",
          "content": "<p>do you discard batches with only negative samples?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961553,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-07T09:26:24.833000",
          "content": "<p>Yep <br>\nI manage to have positive samples in every batch for those expeiments.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961560,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-08-07T09:38:10.430000",
          "content": "<p>cool, out of curiosity do you handle this in pytorch sampler or do you manually discard batches with only zeros?<br>\nWould you share your approach? (after competition end is fine)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961607,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-07T10:30:14.033000",
          "content": "<p>I modified a bit the Pytorch sampler in this <a href=\"https://github.com/ufoym/imbalanced-dataset-sampler\" target=\"_blank\">repo</a> ( in particular the<code>_get_label</code> method to fit the needs here )  . This works for single GPU.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 961922,
          "author_name": "FGPC",
          "author_url": "",
          "post_date": "2020-08-07T16:07:11.473000",
          "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>: you may refer to this discussion regarding data sampling strategy: <a href=\"https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824\" target=\"_blank\">https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation/discussion/107824</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 962545,
          "author_name": "豆柴金鯱",
          "author_url": "",
          "post_date": "2020-08-08T08:14:27.557000",
          "content": "<p>I think this is a right way to do.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1368157,
      "author_name": "nyanp",
      "author_url": "",
      "post_date": "2021-06-28T11:14:28.087000",
      "content": "<p>While class imbalance is not important for auc, it may make a difference for mixup, since most samples will simply be a mix of two negative samples when mixup is performed on imbalanced data.<br>\nSo I just oversampled the positive samples by x3 so that the mixup would try more combinations of positive and negative.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1368246,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-28T12:50:04.227000",
          "content": "<p>I agree, and I do it now as well.  Let me edit the post.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 964057,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-08-09T14:55:51.290000",
      "content": "<p>Imagine if there are 1000000000 negative images and just 1 positive image, then we need to do lots of epochs for back-propagating that positive image effectively.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 964223,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-09T17:09:59.447000",
          "content": "<p>Yes, convergence speed seems the main benefit here.  Experimenting with it.  But have you improved roc-auc with upsampling?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 962174,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-08-07T21:47:23.837000",
      "content": "<p>This seems to be something that is taught in school as massively important. Almost every recent grad I interact with believes this to be very important. My viewpoint is that maybe in some scenarios it can lead to quicker convergence but ultimately what matters is sufficient samples to learn a pattern. if you have 100 positive points and 10,000 negative points your positive points might not be enough to discover the pattern that makes those samples positives, but if you have 10,000 and 10,000,000 positive and negative points you likely have enough positive samples to discover the pattern of those points, but might struggle with convergence. The issue most often is the number of minority points, not the ratio of minority to majority at least with most modern algorithms. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 963092,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2020-08-08T16:58:09.960000",
          "content": "<blockquote>\n  <p>This seems to be something that is taught in school as massively important. Almost every recent grad I interact with believes this to be very important.</p>\n</blockquote>\n<p>Do you think this a hangover from how logistic regression is taught?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 964425,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-08-09T21:00:13.023000",
          "content": "<p>That's my understanding. Traditional methods fall prey to global statistics much worse than non-parametric methods like gradient boosted decision trees. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 962245,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-08T00:27:49.990000",
      "content": "<p>I ran an experiment using triple stratified CV and posted results <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\" target=\"_blank\">here</a>. Using different seeds we consistently get around 0.002 AUC increase with 25x upsample for model RAPIDS cuML kNN on image embeddings. It's small but it helps. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc79785c71cde03a928c41bd2b1cb049d%2Faucc.png?generation=1596846316898566&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 962247,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-08T00:33:23.503000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> just to be clear, you mean you took all malignant images (2019, new, 2020) and added them in 25 times? That's over 100,000 images right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962251,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-08T00:39:11.030000",
          "content": "<p>No. In this plot (experiment), I only use 2020 competition data with my RAPIDS cuML kNN model. After removing duplicates, the 2020 data has 32692 images with 584 malignant for a proportion of 1.79%. Then i add 25 copies of each of 584 malignant (i.e. 14600 images). Now the training data has <code>47292 = 32692 + 25*584</code> images with 32.1% malignant.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 962269,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-08T01:32:25.097000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks for clarifying, and thanks for your post I saw you made on this, very informative.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 962758,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-08T12:18:06.833000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Do I get it right that you increase AUC for knn? I am not surprised, as knn depend son the density of each class. This is an interesting find still, thanks for checking.</p>\n<p>I have yet to see an improvement of AUC for good CNN models. I'd be very happy to see someone posting an improvement.</p>\n<p>So far the only documented improvement is training convergence improved with upsampling. I didn't think of it. It is another good outcome of this discussion for me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 961694,
      "author_name": "Μαριος Μιχαηλιδης KazAnova",
      "author_url": "",
      "post_date": "2020-08-07T12:25:51.210000",
      "content": "<p>I think it depends on the problem. If you know it advance (through probing or other means) that the test dataset might have a different target distribution , then it might worth considering balancing to match that. Otherwise, most of my experiences trying to change the principal target distribution have ended up in failures. </p>\n<p>I have also seen problems when the frequency of the minority class is too small, that balancing (or up sampling that class) in the first few epochs while training NNs  can help the model converge faster.  </p>",
      "votes": 4,
      "replies": [
        {
          "id": 961717,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-07T12:49:58.900000",
          "content": "<blockquote>\n  <p>test dataset might have a different target distribution , then it might worth considering balancing to match that. </p>\n</blockquote>\n<p>How is this relevant when metric is roc-auc?  I agree with you if the metric depend on probabbility calibration.</p>\n<blockquote>\n  <p>balancing (or up sampling that class) in the first few epochs while training NNs can help the model converge faster. </p>\n</blockquote>\n<p>Interesting, thanks for sharing.  I will give it a try if possiible.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 961720,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2020-08-07T12:54:29.727000",
          "content": "<blockquote>\n  <p>How is this relevant when metric is roc-auc? I agree with you if the metric depends on probabbility calibration.</p>\n</blockquote>\n<p>Theoretically, it is not relevant when the metric is AUC. It is more of a general observation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 962515,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2020-08-08T07:29:00.197000",
          "content": "<p>i can confirm the speed up in convergency. thanks for sharing the tip. </p>\n\n<p>do you think there will be benefits in continuous training with the unbalance dataset? reading the conservation above, it seems the answer is no as the metric is auc.  </p>",
          "votes": 0,
          "replies": [
            {
              "id": 962548,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-08-08T08:16:50.450000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 961170,
      "author_name": "seaguII",
      "author_url": "",
      "post_date": "2020-08-07T01:00:09.237000",
      "content": "<p>In my opinion, if the dataset has the same distribution as the real world(for example, your test set), I don't see a necessity to do sampling even though there is class imbalance.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 961455,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-07T07:37:08.430000",
          "content": "<p>That's only when AUC is the metric I hope? In a health care setting, I would care about the positive class ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961592,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-08-07T10:16:53.430000",
          "content": "<p>Gilles, I do not agree. At least not fully. Here is why.</p>\n<p>It is often the case that the most important are actually those negative that turn into false positive. You simply do not wanna falsely indicate presence of a condition such as a disease especially when a treatment also causes harm. Even if not harmful directly, it's still super stressful and should be avoided. <em>Primum non nocere</em></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 961603,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-07T10:24:36.413000",
          "content": "<p>It's mostly much worse the other way around though. Diagnosing a healthy person as sick will cost an unnecessary operation and possibly some harm. Diagnosing a sick person as healthy (e.g. cancer) will cost a human life.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 967324,
          "author_name": "Kishan Joshi",
          "author_url": "",
          "post_date": "2020-08-12T06:57:17.480000",
          "content": "<p>It means we want to focus more on false-negative rather than false-positive.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 967329,
          "author_name": "Kishan Joshi",
          "author_url": "",
          "post_date": "2020-08-12T07:02:33.480000",
          "content": "<p>Ya <a href=\"https://www.kaggle.com/fangao\" target=\"_blank\">@fangao</a> but how then our model will learn minority class from such a small number of samples?<br>\nAre the samples for minority class sufficient for learning?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 961919,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-07T15:54:34.447000",
      "content": "<h3>Confusion between \"upsample\" and \"add data\"</h3>\n<blockquote>\n  <p>There is some confusion about wording however. Upsampling the minority class means that samples form minority class are sampled more than once (duplicated). Obviously. some confuse that with adding new samples for the minority class.</p>\n</blockquote>\n<p>When using TFRecords, many participants \"upsample\" by \"adding more TFRecords\" that contain only malignant images that are already in their train data. This is causing the confusion between the words \"upsample\" and \"add data\". I explain how to \"upsample\" by \"adding data\" <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a> and share TFRecords containing only malignant. </p>\n<blockquote>\n  <p>Adding more data helps in general. Adding only samples from minority class does not help necessarily however. </p>\n</blockquote>\n<p>If we add samples from the minority class that are already present in our train, then we are \"upsampling\". If the samples from the minority class are not already present in our train, we are \"not upsampling\". If we wish to experiment with \"upsampling\" 2019 external data, then we first add my 2019 external dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a> and second add my 2019 malignant dataset <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>. (To \"upsample\" 2020 data, we add my 2020 malignant <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a>).</p>\n<h3>Example \"upsample\" Notebook</h3>\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\" target=\"_blank\">here</a> that adds 2019 data and then \"upsamples\" the 2019 data. It also \"upsamples\" the 2020 data. Both \"upsample\" by \"adding TFRecords\". The notebook uses 128x128 images and EfficientNetB0. It scores 0.910 \"CV\" and 0.916 LB. Note how well the CV and LB align. The \"CV\" score is computed by ensembling the 3 experiments as follows:</p>\n<pre><code>from sklearn.metric import roc_auc_score\noof = df_oof.iloc[:6552,2].values + df_oof.iloc[6552:6552*2,2].values\\\n          + df_oof.iloc[6552*2:,2].values\nprint('Overall CV AUC =', roc_auc_score(df.iloc[:6552,1].values,oof) )\n# Overall CV AUC = 0.909766721673346\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F557caa6c4821d8ed7f3644cadd43766c%2Flb.png?generation=1596815325323718&amp;alt=media\" alt=\"\"> </p>",
      "votes": 3,
      "replies": [
        {
          "id": 961939,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-07T16:25:44.127000",
          "content": "<blockquote>\n  <p>Why Confusion Exists</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Indeed, it probably comes from your posts where you both upsample and add new data.  </p>\n<p>As I wrote, terminology isn't that important as long as we speak about the same.  Given the first comments above where about adding data rather than upsampling, I thought I had to make it clearer.  I hope we are all set now.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 961940,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-08-07T16:26:29.910000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 964238,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-09T17:15:31.043000",
      "content": "<p>I edited the post, given upsampling seems to enable faster convergence.  This alone justifies upsampling IMHO.  Thanks to all who shared this.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 964359,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-09T19:18:09.403000",
          "content": "<p>It must require some retuning because I cannot match the roc-auc I got without upsampling.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 963731,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2020-08-09T08:37:49.087000",
      "content": "<p>In my brief experience, with a more balanced ratio between the classes you can converge in fewer steps</p>",
      "votes": 1,
      "replies": [
        {
          "id": 964006,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-09T14:07:54.663000",
          "content": "<p>indeed, several other people made the same point.  Thanks for adding more confirmation.  I will certainly try at a  point.</p>\n<p>Does it mean you decreased the number of epochs ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 963427,
      "author_name": "Early Bird",
      "author_url": "",
      "post_date": "2020-08-09T03:07:58.640000",
      "content": "<p>Just a thought... if upsampling without Augmentation, a commonsense tells me it's doing no more than just creating duplicates. However, if upsampling with extensive augmentation, wouldn't the problem be a little bit more debatable? As augmentation is adding variance, which is not necessarily a bad thing to have. This also explains why it makes model converges faster as model is seeing more variations of one particular malignant picture at one epoch. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 964007,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-09T14:08:50.757000",
          "content": "<p>This point was made already as well.  Question still remains: do you get better roc-auc that way?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 962977,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-08-08T15:22:57.543000",
      "content": "<p>Thank you for discussing openly, freely and without tickets. Development and the next best level take place through these analyzes. I listen and take notes.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 963211,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-08T18:54:23.613000",
          "content": "<p>Thank you.  Glad to see that some don't mind me asking questions around ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 961667,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-07T11:29:30.040000",
      "content": "<p>I updated the topic to address the main source of disagreement: here upsampling has its usual meaning, where samples from the minority class are replicated.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 961539,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-08-07T09:15:44.373000",
      "content": "<p>Can it help with extreme or specific upsample-TTA to the upsampled data, does it help or status quo.<br>\nWhen upsample some classes, can it increase training/class weight for some samples and maybe cause a new overfitting when trying to avoid it in the first place? <br>\nSolution(?): Maybe one should use both upsamples and downsample together, do little of both, midway solution and with different amount of +/- samples each process in the same CV fold or the next, same data but different class weight every time.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 961445,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2020-08-07T07:33:40.467000",
      "content": "<p>The only reason I am concerned about target balancing here is deep learning and batches (contrary to boosting models that have a broad view of global statistics).</p>\n<p>We have 2% of positive samples, so if because of model size and image size I can’t use a bigger batch size than 8 or 16, I’ll end up training my model with a lot of batches only containing 0 targets and a few of them containing 1 or 2 ones.<br>\nIt feels like this could hurt the gradient descend or at least slow it down.</p>\n<p>That’s why trying to have at least one positive example per batch could seem fair, or giving more importance to ones so that they end up having a word to say when they show up in one batch.</p>\n<p>Here I agree that it might not drastically change the final score, but if you can get the same result with 10 epochs instead of 20 by doing something about class imbalance then you save a lot of time for more experiments.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 961580,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-08-07T10:00:15.297000",
          "content": "<p>Perhaps you can mitigate this with gradient accumulation?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 961955,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-08-07T16:49:22.477000",
          "content": "<p>yes that's one solution as well! I've never used it but definitely something to try!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 961105,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-08-06T23:06:36.070000",
      "content": "<p>How are you to predict \"malignant\" when you don't have many malignant to train from?  For the order to be correct, you must be able to determine benign vs malignant.  You can't possibly determine malignant unless you have enough data to help you do that.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 961120,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-06T23:53:35.833000",
          "content": "<p>You mix using more data, eg external data, and fixing class imbalance I'm afraid.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 961131,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-07T00:03:13.080000",
          "content": "<p>To be clearer maybe: upsampling from the same data does not add new data.  Same for class weights.  I didn't see improvement when doing it, and I am asking if other see improvement.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 961205,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-08-07T02:08:10.307000",
          "content": "<p>What if you use a heavy set of augmentations? Then upsampling “sort of” adds new data, since (almost) all extra images are going to be unique. </p>\n\n<p>In the end, upsampling increases the share of malignant images in every mini batch. In my understanding, this can help to learn patters from malignant images better, since their contribution to the gradients you compute is higher compared to an original batch with a fewer malignant. </p>\n\n<p>That said, my results so far are  mixed.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 961657,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2020-08-07T11:16:09.363000",
              "content": "<blockquote>\n  <p>What if you use a heavy set of augmentations? Then upsampling “sort of” adds new data, since (almost) all extra images are going to be unique. </p>\n</blockquote>\n<p>If you use the same set of augmentations for positive and negative then you don't change class proportion.  This is not upsampling the minority class.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 961353,
          "author_name": "Vishnu Subramanian",
          "author_url": "",
          "post_date": "2020-08-07T05:42:44.753000",
          "content": "<p>I have not seen any improvement by using Upsampling, by upsampling I mean showing the malignant images more number of times in each epoch then they are. But I was expecting it to work since there are few examples in each batch containing malignant the model may not get enough chance to know about it. But from some of my experimentation, it is not helping. I have a few more ideas that I wanted to try. Hope I can make it work. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967340,
          "author_name": "Kishan Joshi",
          "author_url": "",
          "post_date": "2020-08-12T07:09:40.323000",
          "content": "<p>Maybe it is overfitting on the malignant images. And the model is not able to generalise from the same small set of malignant images. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 962796,
      "author_name": "David J. Slate",
      "author_url": "",
      "post_date": "2020-08-08T12:53:00.953000",
      "content": "<p>For what it's worth, this is what I'm currently trying with neural nets, although I can't claim any particular success with it. My nets train, loss drops, and predictions are much better than random chance but very far from medal territory:</p>\n\n<ol>\n<li><p>I'm working on an x86_64 workstation running Linux, R, Keras, Tensorflow, CUDA, and an Nvidia Quadro GP100 GPU.</p></li>\n<li><p>I'm using no external data, just the original JPEG images resized to 384x384. But I also populate some additional input channels besides the 3 RGB with a few other features computed from the images and meta data, each such feature replicating its value across every pixel of its associated channel.</p></li>\n<li><p>I'm using a home-grown neural net of no particular distinction, containing 4 convolutional layers and associated parametric rectified linear activation, batch normalization, and max pooling layers. I flatten the output of those, feed the result to a dense layer with a single output value, and apply a sigmoid activation. I use BCE with some label smoothing as my loss function, and an rmsprop optimizer.</p></li>\n<li><p>I don't train in \"epochs\", but in a long series of individual batches using Keras's \"train_on_batch\" function. Each batch contains 24 samples chosen randomly from the training data, but 3 slots (12.5%) are reserved for positive examples and the remaining 21 for negative. This amounts to a considerable boost in the percentage of positive examples compared to the 1.76% in the training data. To each image in the batch I randomly apply, with specified probabilities, some combination of augmentations chosen currently from \"flip\", \"flop\", and \"transpose\".</p></li>\n<li><p>As training continues it eventually gets stuck making no additional progress reducing the loss. When this happens my algorithm divides the learning rate in half, restores the net's weights from before progress halted, and continues training. After several such restarts (currently 6), if progress halts again training is terminated. This typically happens between 10000 and 15000 batches, which can take several hours to run.</p></li>\n<li><p>I use TTA (test time augmentation) by predicting for all eight combinations of original, flipped, flopped, and/or transposed images, and averaging the results together.</p></li>\n<li><p>To try to reduce overfitting I conduct several runs with different random number seeds and ensemble the results by averaging their predictions after converting them to ranks.</p></li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 963212,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-08T18:55:32.087000",
          "content": "<p>Have you compared with a lower percentage of positives?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 963277,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2020-08-08T20:43:20.353000",
          "content": "<p>Not yet with my current setup, but I may try that, as well as larger %.  I've also been working on other approaches to model-building for this contest, such as lightgbm with a bunch of features computed in R, as well as a rather exotic genetic algorithm specialized for image classification (perhaps to be described in a later post).  So my computers and my own brain have been pretty busy lately.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 962784,
      "author_name": "OMKAR GURAV",
      "author_url": "",
      "post_date": "2020-08-08T12:44:21.673000",
      "content": "<p>While working on a classification problem I noticed that data is slightly imbalanced. So I used data balancing techniques like UpSample and DownSample. \nIn UpSampling data points of minority count is are increased to the level of data points of majority count. It simply creates duplicate data points in the range of same data points. \nIn DowSampling inverse process is done. Some data points from class of majority count are removed in order to match the count of class of minority count.</p>\n\n<p>Find complete work here\n<a href=\"https://www.kaggle.com/omkargurav/loan-approval-prediction\">Loan Approval Prediction</a></p>\n\n<p>`A = list(data.Loan_Status).count(1)\nB = list(data.Loan_Status).count(0)\nprint(\"Count of 1: \",A,\"\\nCount of 0: \",B)</p>\n\n<p>fig = px.bar((A,B),x=[\"Approved\",\"Rejected\"],y=[A,B],color=[A,B])\nfig.show()\nCount of 1:  422 \nCount of 0:  192`</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3559882%2F101b8f6e6a6df24c5ed31afea98318bb%2Fnewplot.png?generation=1596892428969836&amp;alt=media\" alt=\"\"></p>\n\n<p>`#Getting seperated data with 1 and 0 status.\ndf_majority = new_data[new_data.Loan_Status==1]\ndf_minority = new_data[new_data.Loan_Status==0]</p>\n\n<h1>Here we are downsampling the Majority Class Data Points.</h1>\n\n<h1>i.e. We will get equal amount of datapoint as Minority class from Majority class</h1>\n\n<p>df_manjority_downsampled = resample(df_majority,replace=False,n_samples=192,random_state=123)\ndf_downsampled = pd.concat([df_manjority_downsampled,df_minority])\nprint(\"Downsampled data:-&gt;\\n\",df_downsampled.Loan_Status.value_counts())</p>\n\n<h1>Here we are upsampling the Minority Class Data Points.</h1>\n\n<h1>i.e. We will get equal amount of datapoint as Majority class from Minority class</h1>\n\n<p>df_monority_upsampled = resample(df_minority,replace=True,n_samples=422,random_state=123)\ndf_upsampled = pd.concat([df_majority,df_monority_upsampled])\nprint(\"Upsampled data:-&gt;\\n\",df_upsampled.Loan_Status.value_counts())\nDownsampled data:-&gt;\n 1    192\n0    192\nName: Loan_Status, dtype: int64\nUpsampled data:-&gt;\n 1    422\n0    422\nName: Loan_Status, dtype: int64`</p>\n\n<p>After this method accuracy increased by around 8-10%.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 962240,
      "author_name": "Shelly Gabriela Leal",
      "author_url": "",
      "post_date": "2020-08-08T00:21:55.820000",
      "content": "<p>In my case, balancing classes prevented me from getting results with the interference of one majority of data type.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 961727,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-08-07T13:03:02.980000",
      "content": "<p>&gt; I have not found it to be useful here, and my experience is that it is never useful when using deep learning</p>\n\n<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>It does help in some case, like in <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/overview\">Flower Classification with TPUs</a></p>\n\n<p>As shown in my kernel <a href=\"https://www.kaggle.com/yihdarshieh/detailed-guide-to-custom-training-with-tpus#no-oversampling-vs.-oversampling-%24N=300%24\">Detailed guide to custom training with TPUs</a>, oversampling improves on the validation dataset.</p>\n\n<pre><code>without oversampling\n    valid recall: 0.890801499933458\n    valid precision: 0.9074883670834412\n    valid f1: 0.8926506729634966\n\nwith oversampling\n    valid recall: 0.9269954377916817\n    valid precision: 0.9235320148160663\n    valid f1: 0.9227047892043186\n</code></pre>\n\n<p>Even training with oversampling but fewer epochs (so equivalent number of examples trained) is slightly better</p>\n\n<pre><code>    valid recall: 0.9156978000819738\n    valid precision: 0.8914945977622719\n    valid f1: 0.8997503256064424\n</code></pre>\n\n<p>I didn't submit the results to check on the test dataset for the LB score though.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 961736,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-07T13:13:58.153000",
          "content": "<p>In Flower competition, metric is F1 score, which depends on prediction values.  Here the metric is roc-auc, which is independent of prediction values.  roc-auc only depends on the ordering.</p>\n<p>I am not surprised upsampling helps with F1 score, as it leads to better probability calibration.  I hinted at in my topic BTW.  But I have yet to see how upsampling helps for roc-auc.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 961739,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-07T13:16:01.240000",
          "content": "<p>If you would tune F1 thresholds for cutoff, you would get similar boosts. Upsampling - as <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> says - moves the probabilities, so default cutoff might be better.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 961744,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-08-07T13:19:42.160000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I didn't realize that you are talking for roc-auc rather than in general. Sorry to bother.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 961754,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2020-08-07T13:30:14.660000",
              "content": "<p>No need to be sorry, you shared information ;)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 961924,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-08-07T16:07:51.807000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> , do you think that upsampling could only calibrate probabilities but not able to make the model have more corrected predictions?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961950,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-07T16:46:05.983000",
          "content": "<p>In my experience, it's mostly a trade-off between class-specific performances. Most of the times, you sacrifice a bit of performance in the majority class for more performance in the minority class (which is often more relevant). Nevertheless, there are use cases where the global accuracy also increased by oversampling. Moreover, I could (probably) draw artificial examples where, for example. SMOTE would be beneficial for both classes. This is because the biases/assumptions introduced by the over-sampling methods actually helped drawing better decision boundaries. These are, however, very rare cases.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 961970,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2020-08-07T17:04:50.210000",
              "content": "<blockquote>\n  <p>it's mostly a trade-off between class-specific performances.</p>\n</blockquote>\n<p>What does it mean when roc-auc is the metric?  </p>\n<p>Again, I know upsampling helps when the metric depends on the absolute value of predictions.  We are not in this case here.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 961194,
      "author_name": "Marcelo Kittlein",
      "author_url": "",
      "post_date": "2020-08-07T01:44:47.600000",
      "content": "<p>It may not necessarily be related to this thread, but I think the necessary goal for this kind of applications is to be able to correctly classify all malignant cases without erring too much on benign ones. I don't know if the metric used is the best for that ... even if it is for the competition ... will it be the best for the use of the knowledge obtained?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 961130,
      "author_name": "Abhinav Bandari",
      "author_url": "",
      "post_date": "2020-08-07T00:02:02.747000",
      "content": "<p>Class imbalance is a real issue and we should at least try to remedy it. In my experience, class weights work sometimes but aren't reliable. You said you also tried upsampling the positive classes and it didn't work, but it really depends on what specific upsampling methods you used. If you are not convinced, try training on external melanoma images (effectively upsampling the positive class) and see if your scores improve.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 961133,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-07T00:03:43.477000",
          "content": "<blockquote>\n  <p>Class imbalance is a real issue</p>\n</blockquote>\n<p>Why is it a real issue?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 961134,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-07T00:04:41.783000",
          "content": "<blockquote>\n  <p>try training on external melanoma images (effectively upsampling the positive class) and see if your scores improve.</p>\n</blockquote>\n<p>That's not upsampling.  That's using more data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 961140,
          "author_name": "Abhinav Bandari",
          "author_url": "",
          "post_date": "2020-08-07T00:09:08.217000",
          "content": "<p>Yes, sorry I misspoke but it still helps address class imbalance, and there are notebooks here showing that adding external data improved their scores (like Chris Deotte).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 961141,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2020-08-07T00:11:13.280000",
              "content": "<p>Again, using more data helps, no argument here.  I am discussing upsampling, which means that some samples are seen more often than others when training.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 961153,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-07T00:28:49.713000",
          "content": "<p>Just a point.  Literally just about anyone upsampling is also doing augmentation (<em>corrected, I had put TTA</em>).  Which means each upsampled image is unique.......possibly very unique: cutout, saturation, contrast, brightness, rotation, skew, etc, etc, etc.  Whic means, when we talk of upsampling, we are talking about new synthetic data.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 961183,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2020-08-07T01:26:12.253000",
              "content": "<blockquote>\n  <p>Just a point. Literally just about anyone upsampling is also doing TTA. Which means each upsampled image is unique…….possibly very unique: cutout, saturation, contrast, brightness, rotation, skew, etc, etc, etc. Which means, when we talk of upsampling, we are talking about new synthetic data.</p>\n</blockquote>\n<p>TTA is test time augmentation.  At this stage you don't use target values, hence you perform the same TTA for negative and positive samples.  You do not change class imbalance via TTA.</p>\n<p>I wish someone was answering my question instead of discussing other topics.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 967068,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-11T22:15:47.450000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "961043": "It is a genuine question for me.  I have tried a bit upsampling or weighting, and did not get real improvement.\n\nWhy are people so obsessed with the class imbalance?  \n\nI would understand if we care about calibrating predictions.  But we don't. Indeed, roc-auc only depend son the ordering of predictions, not their values.\n\nthe only reason I see is to add diversity in models, which is good when ensembling.  But there are many ways to add diversity, and they are way less discussed than class imbalance.\n\nI'll try again when I will be using larger images and larger models, but so far it looks like focusing on the wrong issue to me.  \n\n**Edit.**  Most comments are very interesting and worth reading.  There is some confusion about wording however.  Upsampling the minority class means that samples form minority class are sampled more than once (duplicated).  Obviously. some confuse that with adding new samples for the minority class. Wording isn't really an issue as long as we discuss the same topic.\n\nAdding more data helps in general. Adding only samples from minority class does not help necessarily however.  I have experienced, and @philippsinger too, that only adding positive samples can hurt actually. \n\nAnyway, let's leave it at this:  adding external data helps in general. \n\nThis post was not about adding external data.  It was about usampling minority class, i.e. stick with the competition data but use Melanoma samples more than the others when training.  I have not found it to be useful here, and my experience is that it is never useful when using deep learning.  I did found it useful with, say XGBoost, in the past.\n\nThere is one reason for upsampling that is worth discussing IMHO, see comments from @Optimo: if batch size is small enough, then many batches have only samples from the majority class.  Some loss functions don't even work in that case, esp the ones approximating auc directly.  My take here is that if you use a loss that works in that case, eg BCE, then this averages out over many batches.  And, as was also pointed to in comments, one can use gradient accumulation to make sure logical batches are large enough.\n\n**Second edit.**  It seems the main benefit, according to several comments, is that training converges faster when upsampling the minority class.  This alone is a good reason for upsampling.  \n\nUnless mistaken no one claimed they got better roc-auc with upsampling for CNN models.  \n\n**Third edit.** Chris Deotte says he ran experiment showing a 0.002 CV AUC improvement with upsampling.  Looking forward to his experiment details.  Anyway, this is the first time someone claims that upsampling improved roc-auc for CNN.  I'll try harder then given I did not found it useful till now.  https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130#964575\n\n**Fourth edit.**  @nyanpn commented that oversampling the minority class can improve mixup, see his comment below.  ",
    "961492": "Doing nothing in imbalanced problems is often the best.",
    "962051": "Small Example I want to add. Maybe related indirectly. Noisy student paper did quite a good experiment on if it is essential to balance data or not when adding external data. In short, they :\n1) Trained Effnet On Imagenet\n2) Predicted Random Images from Interner(out of distribution from original Imagenet Dataset) -&gt; generate Pseudolabels\n3) Added this Images and retrained further. \n\nNow when you predict Images that are not present in the imagenet dataset, you will get different class distribution. The question that was raised is essential to balance classes when retraining network? \n\nExperimental Answer: For small model -Yes! (you will gain slight boost). It was not helpful for bigger models since they have more power to learn and separate features. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2F9a949952781f1d7da2412c95cdc8faf2%2FScreen%20Shot%202020-08-07%20at%202.41.15%20PM.png?generation=1596825750152331&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F991320%2Ff96afd76d6c8301316bdae17d6f020a2%2FScreen%20Shot%202020-08-07%20at%202.41.20%20PM.png?generation=1596825764119842&amp;alt=media)\n\nAs I mentioned its probably very remotely related to the topic but perhaps can give a good experimental evidence for the discussion. \n\nPaper (https://arxiv.org/pdf/1911.04252.pdf)",
    "961490": "I don't upsample but I do some  balancing in some of my experiments for the sole purpose to speed up training. \nFor that, I discard some benign ones , but it's done online i.e when the batch is being fed to the model. Because model don't need to focus that much on benign cases and this save training time too ;)",
    "1368157": "While class imbalance is not important for auc, it may make a difference for mixup, since most samples will simply be a mix of two negative samples when mixup is performed on imbalanced data.\nSo I just oversampled the positive samples by x3 so that the mixup would try more combinations of positive and negative.",
    "964057": "Imagine if there are 1000000000 negative images and just 1 positive image, then we need to do lots of epochs for back-propagating that positive image effectively.",
    "962174": "This seems to be something that is taught in school as massively important. Almost every recent grad I interact with believes this to be very important. My viewpoint is that maybe in some scenarios it can lead to quicker convergence but ultimately what matters is sufficient samples to learn a pattern. if you have 100 positive points and 10,000 negative points your positive points might not be enough to discover the pattern that makes those samples positives, but if you have 10,000 and 10,000,000 positive and negative points you likely have enough positive samples to discover the pattern of those points, but might struggle with convergence. The issue most often is the number of minority points, not the ratio of minority to majority at least with most modern algorithms. ",
    "962245": "I ran an experiment using triple stratified CV and posted results [here][1]. Using different seeds we consistently get around 0.002 AUC increase with 25x upsample for model RAPIDS cuML kNN on image embeddings. It's small but it helps. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc79785c71cde03a928c41bd2b1cb049d%2Faucc.png?generation=1596846316898566&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130",
    "961694": "I think it depends on the problem. If you know it advance (through probing or other means) that the test dataset might have a different target distribution , then it might worth considering balancing to match that. Otherwise, most of my experiences trying to change the principal target distribution have ended up in failures. \n\nI have also seen problems when the frequency of the minority class is too small, that balancing (or up sampling that class) in the first few epochs while training NNs  can help the model converge faster.  ",
    "961170": "In my opinion, if the dataset has the same distribution as the real world(for example, your test set), I don't see a necessity to do sampling even though there is class imbalance.",
    "961919": "### Confusion between \"upsample\" and \"add data\"\n&gt; There is some confusion about wording however. Upsampling the minority class means that samples form minority class are sampled more than once (duplicated). Obviously. some confuse that with adding new samples for the minority class.\n\nWhen using TFRecords, many participants \"upsample\" by \"adding more TFRecords\" that contain only malignant images that are already in their train data. This is causing the confusion between the words \"upsample\" and \"add data\". I explain how to \"upsample\" by \"adding data\" [here][1] and share TFRecords containing only malignant. \n\n&gt; Adding more data helps in general. Adding only samples from minority class does not help necessarily however. \n\nIf we add samples from the minority class that are already present in our train, then we are \"upsampling\". If the samples from the minority class are not already present in our train, we are \"not upsampling\". If we wish to experiment with \"upsampling\" 2019 external data, then we first add my 2019 external dataset [here][2] and second add my 2019 malignant dataset [here][1]. (To \"upsample\" 2020 data, we add my 2020 malignant [here][1]).\n\n### Example \"upsample\" Notebook\n\nI posted a starter notebook [here][3] that adds 2019 data and then \"upsamples\" the 2019 data. It also \"upsamples\" the 2020 data. Both \"upsample\" by \"adding TFRecords\". The notebook uses 128x128 images and EfficientNetB0. It scores 0.910 \"CV\" and 0.916 LB. Note how well the CV and LB align. The \"CV\" score is computed by ensembling the 3 experiments as follows:\n\n    from sklearn.metric import roc_auc_score\n    oof = df_oof.iloc[:6552,2].values + df_oof.iloc[6552:6552*2,2].values\\\n              + df_oof.iloc[6552*2:,2].values\n    print('Overall CV AUC =', roc_auc_score(df.iloc[:6552,1].values,oof) )\n    # Overall CV AUC = 0.909766721673346\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F557caa6c4821d8ed7f3644cadd43766c%2Flb.png?generation=1596815325323718&amp;alt=media) \n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\n[3]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
    "964238": "I edited the post, given upsampling seems to enable faster convergence.  This alone justifies upsampling IMHO.  Thanks to all who shared this.",
    "963731": "In my brief experience, with a more balanced ratio between the classes you can converge in fewer steps",
    "963427": "Just a thought... if upsampling without Augmentation, a commonsense tells me it's doing no more than just creating duplicates. However, if upsampling with extensive augmentation, wouldn't the problem be a little bit more debatable? As augmentation is adding variance, which is not necessarily a bad thing to have. This also explains why it makes model converges faster as model is seeing more variations of one particular malignant picture at one epoch. ",
    "962977": "Thank you for discussing openly, freely and without tickets. Development and the next best level take place through these analyzes. I listen and take notes.",
    "961667": "I updated the topic to address the main source of disagreement: here upsampling has its usual meaning, where samples from the minority class are replicated.",
    "961539": "Can it help with extreme or specific upsample-TTA to the upsampled data, does it help or status quo.\nWhen upsample some classes, can it increase training/class weight for some samples and maybe cause a new overfitting when trying to avoid it in the first place? \nSolution(?): Maybe one should use both upsamples and downsample together, do little of both, midway solution and with different amount of +/- samples each process in the same CV fold or the next, same data but different class weight every time.",
    "961445": "The only reason I am concerned about target balancing here is deep learning and batches (contrary to boosting models that have a broad view of global statistics).\n\nWe have 2% of positive samples, so if because of model size and image size I can’t use a bigger batch size than 8 or 16, I’ll end up training my model with a lot of batches only containing 0 targets and a few of them containing 1 or 2 ones.\nIt feels like this could hurt the gradient descend or at least slow it down.\n\nThat’s why trying to have at least one positive example per batch could seem fair, or giving more importance to ones so that they end up having a word to say when they show up in one batch.\n\nHere I agree that it might not drastically change the final score, but if you can get the same result with 10 epochs instead of 20 by doing something about class imbalance then you save a lot of time for more experiments.",
    "961105": "How are you to predict \"malignant\" when you don't have many malignant to train from?  For the order to be correct, you must be able to determine benign vs malignant.  You can't possibly determine malignant unless you have enough data to help you do that.",
    "962796": "For what it's worth, this is what I'm currently trying with neural nets, although I can't claim any particular success with it. My nets train, loss drops, and predictions are much better than random chance but very far from medal territory:\n\n1. I'm working on an x86_64 workstation running Linux, R, Keras, Tensorflow, CUDA, and an Nvidia Quadro GP100 GPU.\n\n2. I'm using no external data, just the original JPEG images resized to 384x384. But I also populate some additional input channels besides the 3 RGB with a few other features computed from the images and meta data, each such feature replicating its value across every pixel of its associated channel.\n\n3. I'm using a home-grown neural net of no particular distinction, containing 4 convolutional layers and associated parametric rectified linear activation, batch normalization, and max pooling layers. I flatten the output of those, feed the result to a dense layer with a single output value, and apply a sigmoid activation. I use BCE with some label smoothing as my loss function, and an rmsprop optimizer.\n\n4. I don't train in \"epochs\", but in a long series of individual batches using Keras's \"train_on_batch\" function. Each batch contains 24 samples chosen randomly from the training data, but 3 slots (12.5%) are reserved for positive examples and the remaining 21 for negative. This amounts to a considerable boost in the percentage of positive examples compared to the 1.76% in the training data. To each image in the batch I randomly apply, with specified probabilities, some combination of augmentations chosen currently from \"flip\", \"flop\", and \"transpose\".\n\n5. As training continues it eventually gets stuck making no additional progress reducing the loss. When this happens my algorithm divides the learning rate in half, restores the net's weights from before progress halted, and continues training. After several such restarts (currently 6), if progress halts again training is terminated. This typically happens between 10000 and 15000 batches, which can take several hours to run.\n\n6. I use TTA (test time augmentation) by predicting for all eight combinations of original, flipped, flopped, and/or transposed images, and averaging the results together.\n\n7. To try to reduce overfitting I conduct several runs with different random number seeds and ensemble the results by averaging their predictions after converting them to ranks.",
    "962784": "While working on a classification problem I noticed that data is slightly imbalanced. So I used data balancing techniques like UpSample and DownSample. \nIn UpSampling data points of minority count is are increased to the level of data points of majority count. It simply creates duplicate data points in the range of same data points. \nIn DowSampling inverse process is done. Some data points from class of majority count are removed in order to match the count of class of minority count.\n\nFind complete work here\n[Loan Approval Prediction](https://www.kaggle.com/omkargurav/loan-approval-prediction)\n\n`A = list(data.Loan_Status).count(1)\nB = list(data.Loan_Status).count(0)\nprint(\"Count of 1",
    "962240": "In my case, balancing classes prevented me from getting results with the interference of one majority of data type.",
    "961727": "&gt; I have not found it to be useful here, and my experience is that it is never useful when using deep learning\n\n@cpmpml \n\nIt does help in some case, like in [Flower Classification with TPUs](https://www.kaggle.com/c/flower-classification-with-tpus/overview)\n\nAs shown in my kernel [Detailed guide to custom training with TPUs](https://www.kaggle.com/yihdarshieh/detailed-guide-to-custom-training-with-tpus#no-oversampling-vs.-oversampling-$N=300$), oversampling improves on the validation dataset.\n\n    without oversampling\n        valid recall: 0.890801499933458\n        valid precision: 0.9074883670834412\n        valid f1: 0.8926506729634966\n\n    with oversampling\n        valid recall: 0.9269954377916817\n        valid precision: 0.9235320148160663\n        valid f1: 0.9227047892043186\n\nEven training with oversampling but fewer epochs (so equivalent number of examples trained) is slightly better\n\n        valid recall: 0.9156978000819738\n        valid precision: 0.8914945977622719\n        valid f1: 0.8997503256064424\n\nI didn't submit the results to check on the test dataset for the LB score though.",
    "961194": "It may not necessarily be related to this thread, but I think the necessary goal for this kind of applications is to be able to correctly classify all malignant cases without erring too much on benign ones. I don't know if the metric used is the best for that ... even if it is for the competition ... will it be the best for the use of the knowledge obtained?",
    "961130": "Class imbalance is a real issue and we should at least try to remedy it. In my experience, class weights work sometimes but aren't reliable. You said you also tried upsampling the positive classes and it didn't work, but it really depends on what specific upsampling methods you used. If you are not convinced, try training on external melanoma images (effectively upsampling the positive class) and see if your scores improve.",
    "967068": ""
  }
}