{
  "id": 76407,
  "title": "up sampling, balance sampling in a batch",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/76407",
  "author_name": "",
  "post_date": "2019-01-02T14:33:33.387380200Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>interesting, some of my experiments show that balance sampling (equal distribution of different classes) perform worse then just random sample in the public LB.</p>\n\n<p>The local validation could shows different results.</p>\n\n<p>Did anyone has this problem?</p>",
  "messages": [
    {
      "id": "449015",
      "postDate": "01/02/2019 14:33:33",
      "content": "<p>interesting, some of my experiments show that balance sampling (equal distribution of different classes) perform worse then just random sample in the public LB.</p>\n\n<p>The local validation could shows different results.</p>\n\n<p>Did anyone has this problem?</p>",
      "rawMarkdown": "interesting, some of my experiments show that balance sampling (equal distribution of different classes) perform worse then just random sample in the public LB.\n\nThe local validation could shows different results.\n\nDid anyone has this problem?",
      "votes": null
    },
    {
      "id": "449158",
      "postDate": "01/02/2019 18:18:59",
      "content": "<p>I saw exactly that you are saying and ended up with upsampling only rare classes 4-5 times, replicating images ~100 times seems not to be a good idea. More or less equal sampling boosts val score a lot, but public LB goes down. I think the reason for that is that in the train dataset there are many very similar images taken from nearby locations in the same cell (a kind of leak to val). However, test dataset may include images of the same class taken from another cell. It becomes a big problem for rare classes, when only a few images are provided. If you upsample them a lot, val score will improve (but in reality it is just overfitting), while the ability of the model to generalize on images taken from other cells does not improve at all or even degrades. The extended train dataset should be less susaptable to such problem, but I didn't check it explicitly.</p>",
      "rawMarkdown": "I saw exactly that you are saying and ended up with upsampling only rare classes 4-5 times, replicating images ~100 times seems not to be a good idea. More or less equal sampling boosts val score a lot, but public LB goes down. I think the reason for that is that in the train dataset there are many very similar images taken from nearby locations in the same cell (a kind of leak to val). However, test dataset may include images of the same class taken from another cell. It becomes a big problem for rare classes, when only a few images are provided. If you upsample them a lot, val score will improve (but in reality it is just overfitting), while the ability of the model to generalize on images taken from other cells does not improve at all or even degrades. The extended train dataset should be less susaptable to such problem, but I didn't check it explicitly.",
      "votes": null
    },
    {
      "id": "449334",
      "postDate": "01/03/2019 01:38:49",
      "content": "<p>I found that going for an exactly-equal sampling leads to worse results. But I got improvement on the public LB using the log-dampened weights approach suggested by @Tilii here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a></p>",
      "rawMarkdown": "I found that going for an exactly-equal sampling leads to worse results. But I got improvement on the public LB using the log-dampened weights approach suggested by @Tilii here: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065",
      "votes": null
    },
    {
      "id": "449692",
      "postDate": "01/03/2019 16:01:33",
      "content": "<p>Heng - I have learned much from your posts, and have always appreciated your willingness to share.  I made the top 5% in the Nuclei contest and was assisted greatly by your posts.  So thank you!  For this contest, I have had much less time and my scores are not competitive this time around and am only in these last days putting in some serious time to try and move up.  So take the following FWIW!</p>\n\n<p>(a) my val sets are all randomly drawn and balanced based on a kmeans clustering (n=8) of the target labels, \n(b) training was run with kaggle only (\"base\") and kaggle plus oversampling of sparse classes including mixing in hpa data of sparse classes (\"base+oversample\").  I duplicated sparse training examples between 3x-5x depending on count available.</p>\n\n<p>I have tried this experiment with Gap Net, InceptionV3, ResNet18, ResNet50, and InceptionV2Resnet.   In general, \"base\" outperforms base+oversample by ~0.06 for Macro F1, and ~0.03 on public LB </p>\n\n<p>Its worth noting that I have also found that log channel weights on \"base\" yields a 0.03 improvement in public LB.  </p>\n\n<p>Hope that helps!</p>",
      "rawMarkdown": "Heng - I have learned much from your posts, and have always appreciated your willingness to share.  I made the top 5% in the Nuclei contest and was assisted greatly by your posts.  So thank you!  For this contest, I have had much less time and my scores are not competitive this time around and am only in these last days putting in some serious time to try and move up.  So take the following FWIW!\n\n(a) my val sets are all randomly drawn and balanced based on a kmeans clustering (n=8) of the target labels, \n(b) training was run with kaggle only (\"base\") and kaggle plus oversampling of sparse classes including mixing in hpa data of sparse classes (\"base+oversample\").  I duplicated sparse training examples between 3x-5x depending on count available.\n\nI have tried this experiment with Gap Net, InceptionV3, ResNet18, ResNet50, and InceptionV2Resnet.   In general, \"base\" outperforms base+oversample by ~0.06 for Macro F1, and ~0.03 on public LB \n\nIts worth noting that I have also found that log channel weights on \"base\" yields a 0.03 improvement in public LB.  \n\nHope that helps!",
      "votes": null
    },
    {
      "id": "449815",
      "postDate": "01/03/2019 19:39:12",
      "content": "<p>Yes I have observed the similar nature of performance when tried to balance dataset. I believe it's not train oversampling/balancing problem, it's a test distribution problem. Neural networks learn distribution of dataset as well. So if one uses fixed threshold for all classes, the output will closely follow the balanced-train distribution. As we already know from @Iafoss, test distribution closely follows train distribution, only balancing train data without adjusting threshold may not produce optimal results. One ends up detecting more rare classes than actual.</p>",
      "rawMarkdown": "Yes I have observed the similar nature of performance when tried to balance dataset. I believe it's not train oversampling/balancing problem, it's a test distribution problem. Neural networks learn distribution of dataset as well. So if one uses fixed threshold for all classes, the output will closely follow the balanced-train distribution. As we already know from @Iafoss, test distribution closely follows train distribution, only balancing train data without adjusting threshold may not produce optimal results. One ends up detecting more rare classes than actual.",
      "votes": null
    },
    {
      "id": "449951",
      "postDate": "01/04/2019 02:28:48",
      "content": "<p>Heng - would you ever consider neural-style transfer (pytorch has some examples) as a way to generate more data for the rare classes? Might require training a \"style\" for each, possibly using \"noise\" content or maybe same classs images.  Have seen this used for generating artistic styles and wondered about its application here.  If/when you have time to comment sometime would be interested in your views, no hurry.  Always learn a lot from your posts and references and exploring different approaches and possibilities.</p>\n\n<p>with respect to style transfer <br>\n<a href=\"https://arxiv.org/abs/1603.08155\">https://arxiv.org/abs/1603.08155</a>\n<a href=\"https://arxiv.org/pdf/1607.08022.pdf\">https://arxiv.org/pdf/1607.08022.pdf</a></p>\n\n<p>with respect to art  <a href=\"https://arxiv.org/pdf/1706.07068.pdf\">https://arxiv.org/pdf/1706.07068.pdf</a></p>",
      "rawMarkdown": "Heng - would you ever consider neural-style transfer (pytorch has some examples) as a way to generate more data for the rare classes? Might require training a \"style\" for each, possibly using \"noise\" content or maybe same classs images.  Have seen this used for generating artistic styles and wondered about its application here.  If/when you have time to comment sometime would be interested in your views, no hurry.  Always learn a lot from your posts and references and exploring different approaches and possibilities.\n\nwith respect to style transfer  \nhttps://arxiv.org/abs/1603.08155\nhttps://arxiv.org/pdf/1607.08022.pdf\n\nwith respect to art  https://arxiv.org/pdf/1706.07068.pdf",
      "votes": null
    },
    {
      "id": "450112",
      "postDate": "01/04/2019 10:25:48",
      "content": "<p>Probably worth stating which loss function one uses when doing any sampling as this impacts significantly. I've been using variants on focal loss generally but did try BCE with an aggressive balanced sampling strategy once (stable but capped out at about ~0.46 LB).</p>",
      "rawMarkdown": "Probably worth stating which loss function one uses when doing any sampling as this impacts significantly. I've been using variants on focal loss generally but did try BCE with an aggressive balanced sampling strategy once (stable but capped out at about ~0.46 LB).",
      "votes": null
    },
    {
      "id": "451896",
      "postDate": "01/07/2019 21:31:51",
      "content": "<p>I am considering some upsampling at this late date (I have convinced myself at the 60% level that downsampling common classes, which is easier, will not be a good way to go).</p>\n\n<p>For fastai users: is there code within fastai that implements augmented upsampling of rare classes? I have looked on the forums and is seems unlikely. I think I could implement something in time but it would be a project...</p>",
      "rawMarkdown": "I am considering some upsampling at this late date (I have convinced myself at the 60% level that downsampling common classes, which is easier, will not be a good way to go).\n\nFor fastai users: is there code within fastai that implements augmented upsampling of rare classes? I have looked on the forums and is seems unlikely. I think I could implement something in time but it would be a project...",
      "votes": null
    },
    {
      "id": "451915",
      "postDate": "01/07/2019 22:11:25",
      "content": "<p>The simplest way to do it is duplication of particular filenames when you construct a dataset. You can check <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#\">my kernel</a> (Oversampling class) for more details. If you use fast.ai 1.0 and construct everything based on dataframe, u can insert duplicated lines into this dataframe to increase abundance of rare classes.</p>",
      "rawMarkdown": "The simplest way to do it is duplication of particular filenames when you construct a dataset. You can check [my kernel][1] (Oversampling class) for more details. If you use fast.ai 1.0 and construct everything based on dataframe, u can insert duplicated lines into this dataframe to increase abundance of rare classes.\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#",
      "votes": null
    },
    {
      "id": "452009",
      "postDate": "01/08/2019 03:04:07",
      "content": "<p>I use fastai, and I’ve been changing the batch sampler of my training dataset to be pytorch’s BatchSampler wrapped around WeightedRandomSampler. There’s no way I’ve found to pass it at construction right now, so I’m using a hack of setting the batch_sampler member after initialization.</p>",
      "rawMarkdown": "I use fastai, and I’ve been changing the batch sampler of my training dataset to be pytorch’s BatchSampler wrapped around WeightedRandomSampler. There’s no way I’ve found to pass it at construction right now, so I’m using a hack of setting the batch_sampler member after initialization.",
      "votes": null
    },
    {
      "id": "452216",
      "postDate": "01/08/2019 11:15:21",
      "content": "<p>For fastai (not v1) I've been using a (hack) of changing the training names according to a sampling method via a callback at the end of each epoch.</p>",
      "rawMarkdown": "For fastai (not v1) I've been using a (hack) of changing the training names according to a sampling method via a callback at the end of each epoch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 449158,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "01/02/2019 18:18:59",
      "content": "<p>I saw exactly that you are saying and ended up with upsampling only rare classes 4-5 times, replicating images ~100 times seems not to be a good idea. More or less equal sampling boosts val score a lot, but public LB goes down. I think the reason for that is that in the train dataset there are many very similar images taken from nearby locations in the same cell (a kind of leak to val). However, test dataset may include images of the same class taken from another cell. It becomes a big problem for rare classes, when only a few images are provided. If you upsample them a lot, val score will improve (but in reality it is just overfitting), while the ability of the model to generalize on images taken from other cells does not improve at all or even degrades. The extended train dataset should be less susaptable to such problem, but I didn't check it explicitly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 451896,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "01/07/2019 21:31:51",
          "content": "<p>I am considering some upsampling at this late date (I have convinced myself at the 60% level that downsampling common classes, which is easier, will not be a good way to go).</p>\n\n<p>For fastai users: is there code within fastai that implements augmented upsampling of rare classes? I have looked on the forums and is seems unlikely. I think I could implement something in time but it would be a project...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 451915,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/07/2019 22:11:25",
          "content": "<p>The simplest way to do it is duplication of particular filenames when you construct a dataset. You can check <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#\">my kernel</a> (Oversampling class) for more details. If you use fast.ai 1.0 and construct everything based on dataframe, u can insert duplicated lines into this dataframe to increase abundance of rare classes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452009,
          "author_name": "hortonhearsafoo",
          "author_url": "",
          "post_date": "01/08/2019 03:04:07",
          "content": "<p>I use fastai, and I’ve been changing the batch sampler of my training dataset to be pytorch’s BatchSampler wrapped around WeightedRandomSampler. There’s no way I’ve found to pass it at construction right now, so I’m using a hack of setting the batch_sampler member after initialization.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452216,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/08/2019 11:15:21",
          "content": "<p>For fastai (not v1) I've been using a (hack) of changing the training names according to a sampling method via a callback at the end of each epoch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 449334,
      "author_name": "hortonhearsafoo",
      "author_url": "",
      "post_date": "01/03/2019 01:38:49",
      "content": "<p>I found that going for an exactly-equal sampling leads to worse results. But I got improvement on the public LB using the log-dampened weights approach suggested by @Tilii here: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449692,
      "author_name": "jbfarrar",
      "author_url": "",
      "post_date": "01/03/2019 16:01:33",
      "content": "<p>Heng - I have learned much from your posts, and have always appreciated your willingness to share.  I made the top 5% in the Nuclei contest and was assisted greatly by your posts.  So thank you!  For this contest, I have had much less time and my scores are not competitive this time around and am only in these last days putting in some serious time to try and move up.  So take the following FWIW!</p>\n\n<p>(a) my val sets are all randomly drawn and balanced based on a kmeans clustering (n=8) of the target labels, \n(b) training was run with kaggle only (\"base\") and kaggle plus oversampling of sparse classes including mixing in hpa data of sparse classes (\"base+oversample\").  I duplicated sparse training examples between 3x-5x depending on count available.</p>\n\n<p>I have tried this experiment with Gap Net, InceptionV3, ResNet18, ResNet50, and InceptionV2Resnet.   In general, \"base\" outperforms base+oversample by ~0.06 for Macro F1, and ~0.03 on public LB </p>\n\n<p>Its worth noting that I have also found that log channel weights on \"base\" yields a 0.03 improvement in public LB.  </p>\n\n<p>Hope that helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449815,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "01/03/2019 19:39:12",
      "content": "<p>Yes I have observed the similar nature of performance when tried to balance dataset. I believe it's not train oversampling/balancing problem, it's a test distribution problem. Neural networks learn distribution of dataset as well. So if one uses fixed threshold for all classes, the output will closely follow the balanced-train distribution. As we already know from @Iafoss, test distribution closely follows train distribution, only balancing train data without adjusting threshold may not produce optimal results. One ends up detecting more rare classes than actual.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449951,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "01/04/2019 02:28:48",
      "content": "<p>Heng - would you ever consider neural-style transfer (pytorch has some examples) as a way to generate more data for the rare classes? Might require training a \"style\" for each, possibly using \"noise\" content or maybe same classs images.  Have seen this used for generating artistic styles and wondered about its application here.  If/when you have time to comment sometime would be interested in your views, no hurry.  Always learn a lot from your posts and references and exploring different approaches and possibilities.</p>\n\n<p>with respect to style transfer <br>\n<a href=\"https://arxiv.org/abs/1603.08155\">https://arxiv.org/abs/1603.08155</a>\n<a href=\"https://arxiv.org/pdf/1607.08022.pdf\">https://arxiv.org/pdf/1607.08022.pdf</a></p>\n\n<p>with respect to art  <a href=\"https://arxiv.org/pdf/1706.07068.pdf\">https://arxiv.org/pdf/1706.07068.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 450112,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "01/04/2019 10:25:48",
      "content": "<p>Probably worth stating which loss function one uses when doing any sampling as this impacts significantly. I've been using variants on focal loss generally but did try BCE with an aggressive balanced sampling strategy once (stable but capped out at about ~0.46 LB).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "449015": "interesting, some of my experiments show that balance sampling (equal distribution of different classes) perform worse then just random sample in the public LB.\n\nThe local validation could shows different results.\n\nDid anyone has this problem?",
    "449158": "I saw exactly that you are saying and ended up with upsampling only rare classes 4-5 times, replicating images ~100 times seems not to be a good idea. More or less equal sampling boosts val score a lot, but public LB goes down. I think the reason for that is that in the train dataset there are many very similar images taken from nearby locations in the same cell (a kind of leak to val). However, test dataset may include images of the same class taken from another cell. It becomes a big problem for rare classes, when only a few images are provided. If you upsample them a lot, val score will improve (but in reality it is just overfitting), while the ability of the model to generalize on images taken from other cells does not improve at all or even degrades. The extended train dataset should be less susaptable to such problem, but I didn't check it explicitly.",
    "449334": "I found that going for an exactly-equal sampling leads to worse results. But I got improvement on the public LB using the log-dampened weights approach suggested by @Tilii here: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065",
    "449692": "Heng - I have learned much from your posts, and have always appreciated your willingness to share.  I made the top 5% in the Nuclei contest and was assisted greatly by your posts.  So thank you!  For this contest, I have had much less time and my scores are not competitive this time around and am only in these last days putting in some serious time to try and move up.  So take the following FWIW!\n\n(a) my val sets are all randomly drawn and balanced based on a kmeans clustering (n=8) of the target labels, \n(b) training was run with kaggle only (\"base\") and kaggle plus oversampling of sparse classes including mixing in hpa data of sparse classes (\"base+oversample\").  I duplicated sparse training examples between 3x-5x depending on count available.\n\nI have tried this experiment with Gap Net, InceptionV3, ResNet18, ResNet50, and InceptionV2Resnet.   In general, \"base\" outperforms base+oversample by ~0.06 for Macro F1, and ~0.03 on public LB \n\nIts worth noting that I have also found that log channel weights on \"base\" yields a 0.03 improvement in public LB.  \n\nHope that helps!",
    "449815": "Yes I have observed the similar nature of performance when tried to balance dataset. I believe it's not train oversampling/balancing problem, it's a test distribution problem. Neural networks learn distribution of dataset as well. So if one uses fixed threshold for all classes, the output will closely follow the balanced-train distribution. As we already know from @Iafoss, test distribution closely follows train distribution, only balancing train data without adjusting threshold may not produce optimal results. One ends up detecting more rare classes than actual.",
    "449951": "Heng - would you ever consider neural-style transfer (pytorch has some examples) as a way to generate more data for the rare classes? Might require training a \"style\" for each, possibly using \"noise\" content or maybe same classs images.  Have seen this used for generating artistic styles and wondered about its application here.  If/when you have time to comment sometime would be interested in your views, no hurry.  Always learn a lot from your posts and references and exploring different approaches and possibilities.\n\nwith respect to style transfer  \nhttps://arxiv.org/abs/1603.08155\nhttps://arxiv.org/pdf/1607.08022.pdf\n\nwith respect to art  https://arxiv.org/pdf/1706.07068.pdf",
    "450112": "Probably worth stating which loss function one uses when doing any sampling as this impacts significantly. I've been using variants on focal loss generally but did try BCE with an aggressive balanced sampling strategy once (stable but capped out at about ~0.46 LB).",
    "451896": "I am considering some upsampling at this late date (I have convinced myself at the 60% level that downsampling common classes, which is easier, will not be a good way to go).\n\nFor fastai users: is there code within fastai that implements augmented upsampling of rare classes? I have looked on the forums and is seems unlikely. I think I could implement something in time but it would be a project...",
    "451915": "The simplest way to do it is duplication of particular filenames when you construct a dataset. You can check [my kernel][1] (Oversampling class) for more details. If you use fast.ai 1.0 and construct everything based on dataframe, u can insert duplicated lines into this dataframe to increase abundance of rare classes.\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#",
    "452009": "I use fastai, and I’ve been changing the batch sampler of my training dataset to be pytorch’s BatchSampler wrapped around WeightedRandomSampler. There’s no way I’ve found to pass it at construction right now, so I’m using a hack of setting the batch_sampler member after initialization.",
    "452216": "For fastai (not v1) I've been using a (hack) of changing the training names according to a sampling method via a callback at the end of each epoch."
  },
  "source": "meta"
}