{
  "id": 71066,
  "title": "Fighting label imbalance",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/71066",
  "author_name": "",
  "post_date": "2018-11-09T20:54:15.639546300Z",
  "votes": 9,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Now, there are two ways to deal with the problem: weighted loss or smarter batching.</p>\n\n<p>A short snippet of code that you can add before the backprop of each batch to see what class is missing from the batch:</p>\n\n<pre><code>inspection = truth.detach().numpy()\nfrom sklearn.preprocessing import MultiLabelBinarizer\nmlb = MultiLabelBinarizer()\nmlb.fit([list(range(28))])\ntruthl = mlb.inverse_transform(inspection)\n\n for i in range(28):\n            logic = False\n            for at in truthl:\n                if i in at: logic=True\n            if not(logic): print(i)\n print(\"====\")\n</code></pre>",
  "messages": [
    {
      "id": "418395",
      "postDate": "11/09/2018 20:54:15",
      "content": "<p>Now, there are two ways to deal with the problem: weighted loss or smarter batching.</p>\n\n<p>A short snippet of code that you can add before the backprop of each batch to see what class is missing from the batch:</p>\n\n<pre><code>inspection = truth.detach().numpy()\nfrom sklearn.preprocessing import MultiLabelBinarizer\nmlb = MultiLabelBinarizer()\nmlb.fit([list(range(28))])\ntruthl = mlb.inverse_transform(inspection)\n\n for i in range(28):\n            logic = False\n            for at in truthl:\n                if i in at: logic=True\n            if not(logic): print(i)\n print(\"====\")\n</code></pre>",
      "rawMarkdown": "Now, there are two ways to deal with the problem: weighted loss or smarter batching.\n\nA short snippet of code that you can add before the backprop of each batch to see what class is missing from the batch:\n\n    inspection = truth.detach().numpy()\n    from sklearn.preprocessing import MultiLabelBinarizer\n    mlb = MultiLabelBinarizer()\n    mlb.fit([list(range(28))])\n    truthl = mlb.inverse_transform(inspection)\n\n     for i in range(28):\n                logic = False\n                for at in truthl:\n                    if i in at: logic=True\n                if not(logic): print(i)\n     print(\"====\")",
      "votes": null
    },
    {
      "id": "418397",
      "postDate": "11/09/2018 21:07:50",
      "content": "<p>Option 3. A different CV scheme</p>",
      "rawMarkdown": "Option 3. A different CV scheme",
      "votes": null
    },
    {
      "id": "418484",
      "postDate": "11/10/2018 01:01:14",
      "content": "<p>A PR curve would be more helpful than printing out a small amount of data. TensorboradX gives you counts of TP, TN, FP, TN as well as a nice curve.</p>\n\n<p><img src=\"https://image.ibb.co/gJJSxq/unnamed.png\" alt=\"enter code here\"></p>\n\n<pre><code>from tensorboardX import SummaryWriter\nwriter = SummaryWriter(config.DIRECTORY_CHECKPOINT)\nprint(\"Tensorboard: \" + \"python .local/lib/python2.7/site-packages/tensorboard/main.py --logdir=RedstoneTorch/\" + config.DIRECTORY_CHECKPOINT + \" --port=6006\")\ndef write_pr_curve(writer, label, predicted, epoch, fold):\n    writer.add_pr_curve(\"eval/pr_curve/{}\".format(fold), label, predicted, epoch)\n</code></pre>",
      "rawMarkdown": "A PR curve would be more helpful than printing out a small amount of data. TensorboradX gives you counts of TP, TN, FP, TN as well as a nice curve.\n\n![enter code here][1]\n\n    from tensorboardX import SummaryWriter\n    writer = SummaryWriter(config.DIRECTORY_CHECKPOINT)\n    print(\"Tensorboard: \" + \"python .local/lib/python2.7/site-packages/tensorboard/main.py --logdir=RedstoneTorch/\" + config.DIRECTORY_CHECKPOINT + \" --port=6006\")\n    def write_pr_curve(writer, label, predicted, epoch, fold):\n        writer.add_pr_curve(\"eval/pr_curve/{}\".format(fold), label, predicted, epoch)\n\n\n  [1]: https://image.ibb.co/gJJSxq/unnamed.png",
      "votes": null
    },
    {
      "id": "418505",
      "postDate": "11/10/2018 01:56:23",
      "content": "<p>The point I was trying to make was that it is necessary to have at least one example per class in a batch. Otherwise with RandomSampler the model will only focus on optimizing majority classes, giving a false sense of \"overfitting.\" Without weighted BCE or balanced sampling training accuracy can easily reach 100% while still having a rather low validation accuracy, even with focal loss gamma=5.</p>",
      "rawMarkdown": "The point I was trying to make was that it is necessary to have at least one example per class in a batch. Otherwise with RandomSampler the model will only focus on optimizing majority classes, giving a false sense of \"overfitting.\" Without weighted BCE or balanced sampling training accuracy can easily reach 100% while still having a rather low validation accuracy, even with focal loss gamma=5.",
      "votes": null
    },
    {
      "id": "418507",
      "postDate": "11/10/2018 02:00:56",
      "content": "<p>I'm using 5-fold random split. So far seems that with the naive classifier approach all models, even top performing ones, are suffering from the huge gap between train and validation loss (you may refer to @Brian 's posts). And due to how macro F1 is calculated, CV might not be trustworthy if not done correctly: e.g. a 10-fold CV, for rods and rings there will be only 1 sample in the validation set that contributes 1/28~0.0357 to the metric. I might be considering finding external data sources, hand-label them (with help from domain knowledge experts), and add them to the training/validation set.</p>",
      "rawMarkdown": "I'm using 5-fold random split. So far seems that with the naive classifier approach all models, even top performing ones, are suffering from the huge gap between train and validation loss (you may refer to @Brian 's posts). And due to how macro F1 is calculated, CV might not be trustworthy if not done correctly: e.g. a 10-fold CV, for rods and rings there will be only 1 sample in the validation set that contributes 1/28~0.0357 to the metric. I might be considering finding external data sources, hand-label them (with help from domain knowledge experts), and add them to the training/validation set.",
      "votes": null
    },
    {
      "id": "418512",
      "postDate": "11/10/2018 02:14:57",
      "content": "<p>For a random sampler, a batch would mostly consist of the five majority classes, and a loss without weighted balancing will heavily favor those, which is not a desired feature. Using weighted loss might be fine, but I suspect that since the label imbalance in the existing training set is as high as 1000:1, the weight on the minority classes would be too high so that it might destabilize training. Another possible alternative is weighted random sampler, where the probability of each sample being chosen is adjusted by the class frequency in the training set. When I tested it out, however, this approach is still more or less unstable where occasionally the batch is missing 2-3 classes (tested by the code snippet above). So an alternative method is needed. I would also note that with the above approaches, heavy augmentations are needed, or the model would choose to \"memorize\" the minority samples.</p>",
      "rawMarkdown": "For a random sampler, a batch would mostly consist of the five majority classes, and a loss without weighted balancing will heavily favor those, which is not a desired feature. Using weighted loss might be fine, but I suspect that since the label imbalance in the existing training set is as high as 1000:1, the weight on the minority classes would be too high so that it might destabilize training. Another possible alternative is weighted random sampler, where the probability of each sample being chosen is adjusted by the class frequency in the training set. When I tested it out, however, this approach is still more or less unstable where occasionally the batch is missing 2-3 classes (tested by the code snippet above). So an alternative method is needed. I would also note that with the above approaches, heavy augmentations are needed, or the model would choose to \"memorize\" the minority samples.",
      "votes": null
    },
    {
      "id": "418585",
      "postDate": "11/10/2018 07:15:07",
      "content": "<p>Perhaps a related questiin: there are some strong correlations in the data (e.g. label 6 is with 25 90% of the time). As a newbie, i ask whether the model will automatically fulfill those correlations or whether they should be encouraged somehow.</p>",
      "rawMarkdown": "Perhaps a related questiin: there are some strong correlations in the data (e.g. label 6 is with 25 90% of the time). As a newbie, i ask whether the model will automatically fulfill those correlations or whether they should be encouraged somehow.",
      "votes": null
    },
    {
      "id": "418661",
      "postDate": "11/10/2018 11:34:56",
      "content": "<p><a href=\"/alexanderliao\">@alexanderliao</a> I agree with this. I have tried using variations on weighting the loss but found it a little unstable too. I do have another idea though:</p>\n\n<p>Step 1: train your model as normal on each fold. We now have a model that is reasonably good at recognising (some) types of proteins.</p>\n\n<p>Step 2: fine-tune your model on massively down-sampled training data to better balance the class distribution. (Obviously you still need to validate on your usual full validation set with all the class imbalance).</p>\n\n<p>This is simple enough if you want to do it and isn't too dissimilar to your idea. </p>\n\n<p>Now, the hard part to get right is clearly step 2 concerning two main questions: how much do you decide to balance your classes and how big should your down-sampled training set be?</p>\n\n<p>The idea is to measure the class imbalance by calculating the entropy of the training distribution and then using Bayesian Optimization (e. g. Gaussian Processes or similar) to find the best down-sampling strategy.</p>\n\n<p>Pseudo-algorithm for each run (not sure yet if this can be per epoch or not):</p>\n\n<pre><code> 1. Record the size of the training data and entropy of the training distribution and run model\n 2. Calculate the validation score\n 3. Run Bayesian Optimization to suggest next training data size and target entropy\n 4. Sample training data on the outcome of step 4\n 5. Go to step 1\n</code></pre>\n\n<p>Disclaimer: I literally only thought of this about 10 minutes ago so have no idea if it will work (or if I've missed something silly that means it's not possible). Keen to hear what people think if they try it.</p>",
      "rawMarkdown": "alexanderliao I agree with this. I have tried using variations on weighting the loss but found it a little unstable too. I do have another idea though:\n\nStep 1: train your model as normal on each fold. We now have a model that is reasonably good at recognising (some) types of proteins.\n\nStep 2: fine-tune your model on massively down-sampled training data to better balance the class distribution. (Obviously you still need to validate on your usual full validation set with all the class imbalance).\n\nThis is simple enough if you want to do it and isn't too dissimilar to your idea. \n\nNow, the hard part to get right is clearly step 2 concerning two main questions: how much do you decide to balance your classes and how big should your down-sampled training set be?\n\nThe idea is to measure the class imbalance by calculating the entropy of the training distribution and then using Bayesian Optimization (e. g. Gaussian Processes or similar) to find the best down-sampling strategy.\n\nPseudo-algorithm for each run (not sure yet if this can be per epoch or not):\n\n     1. Record the size of the training data and entropy of the training distribution and run model\n     2. Calculate the validation score\n     3. Run Bayesian Optimization to suggest next training data size and target entropy\n     4. Sample training data on the outcome of step 4\n     5. Go to step 1\n\nDisclaimer: I literally only thought of this about 10 minutes ago so have no idea if it will work (or if I've missed something silly that means it's not possible). Keen to hear what people think if they try it.",
      "votes": null
    },
    {
      "id": "418687",
      "postDate": "11/10/2018 12:32:54",
      "content": "<p>It probably won't do it perfectly, so some post-processing is definitely required.</p>",
      "rawMarkdown": "It probably won't do it perfectly, so some post-processing is definitely required.",
      "votes": null
    },
    {
      "id": "419016",
      "postDate": "11/11/2018 05:10:32",
      "content": "<p>Paper: <a href=\"https://www.cs.umd.edu/~emhand/Papers/AAAI2018_SelectiveLearning.pdf\">https://www.cs.umd.edu/~emhand/Papers/AAAI2018_SelectiveLearning.pdf</a>\n<img src=\"https://i.postimg.cc/NfcYzBV0/1.png\" alt=\"from csdn\">\nBalancing the labels within the batch by:\n  1. Dropout over-represented labels\n  2. Repeat under-represented</p>\n\n<p>But I don't see any benefits compared to weighted BCE.\nIt just that you can use any loss function.\nSo the best method probably still is: (if you don't want to train 28 different model)\n  1. Increase batch size\n  2. repeat under-represented labels from outside of the batch\n  3. output accuracy for each class to see overfitting\n  4. pseudo-labeling</p>\n\n<p><img src=\"https://i.postimg.cc/pV8RCps7/2.png\" alt=\"from csdn\">\n<img src=\"https://i.postimg.cc/YqzkcTKz/3.png\" alt=\"from csdn\">\n<img src=\"https://i.postimg.cc/FzjvHBfm/4.png\" alt=\"from csdn\"></p>",
      "rawMarkdown": "Paper: https://www.cs.umd.edu/~emhand/Papers/AAAI2018_SelectiveLearning.pdf\n![from csdn][1]\nBalancing the labels within the batch by:\n  1. Dropout over-represented labels\n  2. Repeat under-represented\n\nBut I don't see any benefits compared to weighted BCE.\nIt just that you can use any loss function.\nSo the best method probably still is: (if you don't want to train 28 different model)\n  1. Increase batch size\n  2. repeat under-represented labels from outside of the batch\n  3. output accuracy for each class to see overfitting\n  4. pseudo-labeling\n\n![from csdn][2]\n![from csdn][3]\n![from csdn][4]\n\n  [1]: https://i.postimg.cc/NfcYzBV0/1.png\n  [2]: https://i.postimg.cc/pV8RCps7/2.png\n  [3]: https://i.postimg.cc/YqzkcTKz/3.png\n  [4]: https://i.postimg.cc/FzjvHBfm/4.png",
      "votes": null
    },
    {
      "id": "419031",
      "postDate": "11/11/2018 05:41:58",
      "content": "<p>Note: The label imbalance is at a scale of 1:1000.</p>",
      "rawMarkdown": "Note: The label imbalance is at a scale of 1:1000.",
      "votes": null
    },
    {
      "id": "419045",
      "postDate": "11/11/2018 06:10:32",
      "content": "<p>I do a random dropout on the high labels, removing 60% of the values 0 and 25.</p>",
      "rawMarkdown": "I do a random dropout on the high labels, removing 60% of the values 0 and 25.",
      "votes": null
    },
    {
      "id": "423294",
      "postDate": "11/17/2018 21:36:00",
      "content": "<p>In my personal opinion, the model's automatical fulfillment of those correlations would likely result in a disaster: classifying a <code>label=ocean</code> image as <code>prediction=boat</code>.</p>",
      "rawMarkdown": "In my personal opinion, the model's automatical fulfillment of those correlations would likely result in a disaster: classifying a `label=ocean` image as `prediction=boat`.",
      "votes": null
    },
    {
      "id": "423717",
      "postDate": "11/18/2018 23:13:19",
      "content": "<p>I use the plotly code to generate a heatmap as show in this kernel: <a href=\"https://www.kaggle.com/shubhammank/atlas-image-classification-edm-starter\">https://www.kaggle.com/shubhammank/atlas-image-classification-edm-starter</a></p>",
      "rawMarkdown": "I use the plotly code to generate a heatmap as show in this kernel: https://www.kaggle.com/shubhammank/atlas-image-classification-edm-starter",
      "votes": null
    },
    {
      "id": "428634",
      "postDate": "11/27/2018 16:18:30",
      "content": "<ol>\n<li><p>Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling</p></li>\n<li><p>Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning</p></li>\n</ol>\n\n<p>Choice of solution:\nThe thumb rule would be to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling.</p>",
      "rawMarkdown": "1. Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling\n\n2. Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning\n\nChoice of solution:\nThe thumb rule would be to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 418397,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "11/09/2018 21:07:50",
      "content": "<p>Option 3. A different CV scheme</p>",
      "votes": null,
      "replies": [
        {
          "id": 418507,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "11/10/2018 02:00:56",
          "content": "<p>I'm using 5-fold random split. So far seems that with the naive classifier approach all models, even top performing ones, are suffering from the huge gap between train and validation loss (you may refer to @Brian 's posts). And due to how macro F1 is calculated, CV might not be trustworthy if not done correctly: e.g. a 10-fold CV, for rods and rings there will be only 1 sample in the validation set that contributes 1/28~0.0357 to the metric. I might be considering finding external data sources, hand-label them (with help from domain knowledge experts), and add them to the training/validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418484,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "11/10/2018 01:01:14",
      "content": "<p>A PR curve would be more helpful than printing out a small amount of data. TensorboradX gives you counts of TP, TN, FP, TN as well as a nice curve.</p>\n\n<p><img src=\"https://image.ibb.co/gJJSxq/unnamed.png\" alt=\"enter code here\"></p>\n\n<pre><code>from tensorboardX import SummaryWriter\nwriter = SummaryWriter(config.DIRECTORY_CHECKPOINT)\nprint(\"Tensorboard: \" + \"python .local/lib/python2.7/site-packages/tensorboard/main.py --logdir=RedstoneTorch/\" + config.DIRECTORY_CHECKPOINT + \" --port=6006\")\ndef write_pr_curve(writer, label, predicted, epoch, fold):\n    writer.add_pr_curve(\"eval/pr_curve/{}\".format(fold), label, predicted, epoch)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 418505,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "11/10/2018 01:56:23",
          "content": "<p>The point I was trying to make was that it is necessary to have at least one example per class in a batch. Otherwise with RandomSampler the model will only focus on optimizing majority classes, giving a false sense of \"overfitting.\" Without weighted BCE or balanced sampling training accuracy can easily reach 100% while still having a rather low validation accuracy, even with focal loss gamma=5.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418512,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "11/10/2018 02:14:57",
      "content": "<p>For a random sampler, a batch would mostly consist of the five majority classes, and a loss without weighted balancing will heavily favor those, which is not a desired feature. Using weighted loss might be fine, but I suspect that since the label imbalance in the existing training set is as high as 1000:1, the weight on the minority classes would be too high so that it might destabilize training. Another possible alternative is weighted random sampler, where the probability of each sample being chosen is adjusted by the class frequency in the training set. When I tested it out, however, this approach is still more or less unstable where occasionally the batch is missing 2-3 classes (tested by the code snippet above). So an alternative method is needed. I would also note that with the above approaches, heavy augmentations are needed, or the model would choose to \"memorize\" the minority samples.</p>",
      "votes": null,
      "replies": [
        {
          "id": 418661,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "11/10/2018 11:34:56",
          "content": "<p><a href=\"/alexanderliao\">@alexanderliao</a> I agree with this. I have tried using variations on weighting the loss but found it a little unstable too. I do have another idea though:</p>\n\n<p>Step 1: train your model as normal on each fold. We now have a model that is reasonably good at recognising (some) types of proteins.</p>\n\n<p>Step 2: fine-tune your model on massively down-sampled training data to better balance the class distribution. (Obviously you still need to validate on your usual full validation set with all the class imbalance).</p>\n\n<p>This is simple enough if you want to do it and isn't too dissimilar to your idea. </p>\n\n<p>Now, the hard part to get right is clearly step 2 concerning two main questions: how much do you decide to balance your classes and how big should your down-sampled training set be?</p>\n\n<p>The idea is to measure the class imbalance by calculating the entropy of the training distribution and then using Bayesian Optimization (e. g. Gaussian Processes or similar) to find the best down-sampling strategy.</p>\n\n<p>Pseudo-algorithm for each run (not sure yet if this can be per epoch or not):</p>\n\n<pre><code> 1. Record the size of the training data and entropy of the training distribution and run model\n 2. Calculate the validation score\n 3. Run Bayesian Optimization to suggest next training data size and target entropy\n 4. Sample training data on the outcome of step 4\n 5. Go to step 1\n</code></pre>\n\n<p>Disclaimer: I literally only thought of this about 10 minutes ago so have no idea if it will work (or if I've missed something silly that means it's not possible). Keen to hear what people think if they try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419016,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/11/2018 05:10:32",
          "content": "<p>Paper: <a href=\"https://www.cs.umd.edu/~emhand/Papers/AAAI2018_SelectiveLearning.pdf\">https://www.cs.umd.edu/~emhand/Papers/AAAI2018_SelectiveLearning.pdf</a>\n<img src=\"https://i.postimg.cc/NfcYzBV0/1.png\" alt=\"from csdn\">\nBalancing the labels within the batch by:\n  1. Dropout over-represented labels\n  2. Repeat under-represented</p>\n\n<p>But I don't see any benefits compared to weighted BCE.\nIt just that you can use any loss function.\nSo the best method probably still is: (if you don't want to train 28 different model)\n  1. Increase batch size\n  2. repeat under-represented labels from outside of the batch\n  3. output accuracy for each class to see overfitting\n  4. pseudo-labeling</p>\n\n<p><img src=\"https://i.postimg.cc/pV8RCps7/2.png\" alt=\"from csdn\">\n<img src=\"https://i.postimg.cc/YqzkcTKz/3.png\" alt=\"from csdn\">\n<img src=\"https://i.postimg.cc/FzjvHBfm/4.png\" alt=\"from csdn\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419031,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "11/11/2018 05:41:58",
          "content": "<p>Note: The label imbalance is at a scale of 1:1000.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419045,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/11/2018 06:10:32",
          "content": "<p>I do a random dropout on the high labels, removing 60% of the values 0 and 25.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418585,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "11/10/2018 07:15:07",
      "content": "<p>Perhaps a related questiin: there are some strong correlations in the data (e.g. label 6 is with 25 90% of the time). As a newbie, i ask whether the model will automatically fulfill those correlations or whether they should be encouraged somehow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 418687,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "11/10/2018 12:32:54",
          "content": "<p>It probably won't do it perfectly, so some post-processing is definitely required.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423294,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/17/2018 21:36:00",
          "content": "<p>In my personal opinion, the model's automatical fulfillment of those correlations would likely result in a disaster: classifying a <code>label=ocean</code> image as <code>prediction=boat</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423717,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/18/2018 23:13:19",
          "content": "<p>I use the plotly code to generate a heatmap as show in this kernel: <a href=\"https://www.kaggle.com/shubhammank/atlas-image-classification-edm-starter\">https://www.kaggle.com/shubhammank/atlas-image-classification-edm-starter</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 428634,
      "author_name": "jman278",
      "author_url": "",
      "post_date": "11/27/2018 16:18:30",
      "content": "<ol>\n<li><p>Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling</p></li>\n<li><p>Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning</p></li>\n</ol>\n\n<p>Choice of solution:\nThe thumb rule would be to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "418395": "Now, there are two ways to deal with the problem: weighted loss or smarter batching.\n\nA short snippet of code that you can add before the backprop of each batch to see what class is missing from the batch:\n\n    inspection = truth.detach().numpy()\n    from sklearn.preprocessing import MultiLabelBinarizer\n    mlb = MultiLabelBinarizer()\n    mlb.fit([list(range(28))])\n    truthl = mlb.inverse_transform(inspection)\n\n     for i in range(28):\n                logic = False\n                for at in truthl:\n                    if i in at: logic=True\n                if not(logic): print(i)\n     print(\"====\")",
    "418397": "Option 3. A different CV scheme",
    "418484": "A PR curve would be more helpful than printing out a small amount of data. TensorboradX gives you counts of TP, TN, FP, TN as well as a nice curve.\n\n![enter code here][1]\n\n    from tensorboardX import SummaryWriter\n    writer = SummaryWriter(config.DIRECTORY_CHECKPOINT)\n    print(\"Tensorboard: \" + \"python .local/lib/python2.7/site-packages/tensorboard/main.py --logdir=RedstoneTorch/\" + config.DIRECTORY_CHECKPOINT + \" --port=6006\")\n    def write_pr_curve(writer, label, predicted, epoch, fold):\n        writer.add_pr_curve(\"eval/pr_curve/{}\".format(fold), label, predicted, epoch)\n\n\n  [1]: https://image.ibb.co/gJJSxq/unnamed.png",
    "418505": "The point I was trying to make was that it is necessary to have at least one example per class in a batch. Otherwise with RandomSampler the model will only focus on optimizing majority classes, giving a false sense of \"overfitting.\" Without weighted BCE or balanced sampling training accuracy can easily reach 100% while still having a rather low validation accuracy, even with focal loss gamma=5.",
    "418507": "I'm using 5-fold random split. So far seems that with the naive classifier approach all models, even top performing ones, are suffering from the huge gap between train and validation loss (you may refer to @Brian 's posts). And due to how macro F1 is calculated, CV might not be trustworthy if not done correctly: e.g. a 10-fold CV, for rods and rings there will be only 1 sample in the validation set that contributes 1/28~0.0357 to the metric. I might be considering finding external data sources, hand-label them (with help from domain knowledge experts), and add them to the training/validation set.",
    "418512": "For a random sampler, a batch would mostly consist of the five majority classes, and a loss without weighted balancing will heavily favor those, which is not a desired feature. Using weighted loss might be fine, but I suspect that since the label imbalance in the existing training set is as high as 1000:1, the weight on the minority classes would be too high so that it might destabilize training. Another possible alternative is weighted random sampler, where the probability of each sample being chosen is adjusted by the class frequency in the training set. When I tested it out, however, this approach is still more or less unstable where occasionally the batch is missing 2-3 classes (tested by the code snippet above). So an alternative method is needed. I would also note that with the above approaches, heavy augmentations are needed, or the model would choose to \"memorize\" the minority samples.",
    "418585": "Perhaps a related questiin: there are some strong correlations in the data (e.g. label 6 is with 25 90% of the time). As a newbie, i ask whether the model will automatically fulfill those correlations or whether they should be encouraged somehow.",
    "418661": "alexanderliao I agree with this. I have tried using variations on weighting the loss but found it a little unstable too. I do have another idea though:\n\nStep 1: train your model as normal on each fold. We now have a model that is reasonably good at recognising (some) types of proteins.\n\nStep 2: fine-tune your model on massively down-sampled training data to better balance the class distribution. (Obviously you still need to validate on your usual full validation set with all the class imbalance).\n\nThis is simple enough if you want to do it and isn't too dissimilar to your idea. \n\nNow, the hard part to get right is clearly step 2 concerning two main questions: how much do you decide to balance your classes and how big should your down-sampled training set be?\n\nThe idea is to measure the class imbalance by calculating the entropy of the training distribution and then using Bayesian Optimization (e. g. Gaussian Processes or similar) to find the best down-sampling strategy.\n\nPseudo-algorithm for each run (not sure yet if this can be per epoch or not):\n\n     1. Record the size of the training data and entropy of the training distribution and run model\n     2. Calculate the validation score\n     3. Run Bayesian Optimization to suggest next training data size and target entropy\n     4. Sample training data on the outcome of step 4\n     5. Go to step 1\n\nDisclaimer: I literally only thought of this about 10 minutes ago so have no idea if it will work (or if I've missed something silly that means it's not possible). Keen to hear what people think if they try it.",
    "418687": "It probably won't do it perfectly, so some post-processing is definitely required.",
    "419016": "Paper: https://www.cs.umd.edu/~emhand/Papers/AAAI2018_SelectiveLearning.pdf\n![from csdn][1]\nBalancing the labels within the batch by:\n  1. Dropout over-represented labels\n  2. Repeat under-represented\n\nBut I don't see any benefits compared to weighted BCE.\nIt just that you can use any loss function.\nSo the best method probably still is: (if you don't want to train 28 different model)\n  1. Increase batch size\n  2. repeat under-represented labels from outside of the batch\n  3. output accuracy for each class to see overfitting\n  4. pseudo-labeling\n\n![from csdn][2]\n![from csdn][3]\n![from csdn][4]\n\n  [1]: https://i.postimg.cc/NfcYzBV0/1.png\n  [2]: https://i.postimg.cc/pV8RCps7/2.png\n  [3]: https://i.postimg.cc/YqzkcTKz/3.png\n  [4]: https://i.postimg.cc/FzjvHBfm/4.png",
    "419031": "Note: The label imbalance is at a scale of 1:1000.",
    "419045": "I do a random dropout on the high labels, removing 60% of the values 0 and 25.",
    "423294": "In my personal opinion, the model's automatical fulfillment of those correlations would likely result in a disaster: classifying a `label=ocean` image as `prediction=boat`.",
    "423717": "I use the plotly code to generate a heatmap as show in this kernel: https://www.kaggle.com/shubhammank/atlas-image-classification-edm-starter",
    "428634": "1. Sampling Solutions:\na. Oversampling:\ni. Random Oversampling: Create duplicates of minority class rows [Might lead to overfitting]\nii. Informed Oversampling: e.g. SMOTE\nb. Undersampling:\ni. Random Undersampling: Randomly remove rows of majority class [Might lose out on important data points]\nii. Informed Undersampling: e.g. Tomek Undersampling\n\n2. Cost-Sensitive Learning:\nAssociates a different cost to false positives and false negatives while training the model.\nDifferent cost combinations will have to be tried using grid search to arrive at the most optimal solution.\nrpart function in R has a feature that allows for cost-sensitive learning\n\nChoice of solution:\nThe thumb rule would be to use cost-sensitive learning wherever possible but not every algorithm allows the user to do so. In such cases go with oversampling/undersampling."
  },
  "source": "meta"
}