{
  "id": 41523,
  "title": "how to use hard samples?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/41523",
  "author_name": "",
  "post_date": "2017-10-19T16:02:23.446605300Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>based on the prediction scores, it can identify which training (and testing) samples are difficult. For example, now i have: {x_i,y_i,s_i} for i=1 to N train samples. x is the image, y is the ground truth label, s is real value to indicate the difficulty. </p>\n\n<p>How can i use use the value of s to improve my training?</p>\n\n<p>E.g. how to formulate better loss?</p>\n\n<p>E.g. how to do better sampling?</p>\n\n<p>This is related to Self-paced Learning or Curriculum Learning. Although I have read about such methods, I have never tried them before. If you can used them before, please give me some advice. Thanks!</p>",
  "messages": [
    {
      "id": "233214",
      "postDate": "10/19/2017 16:02:23",
      "content": "<p>based on the prediction scores, it can identify which training (and testing) samples are difficult. For example, now i have: {x_i,y_i,s_i} for i=1 to N train samples. x is the image, y is the ground truth label, s is real value to indicate the difficulty. </p>\n\n<p>How can i use use the value of s to improve my training?</p>\n\n<p>E.g. how to formulate better loss?</p>\n\n<p>E.g. how to do better sampling?</p>\n\n<p>This is related to Self-paced Learning or Curriculum Learning. Although I have read about such methods, I have never tried them before. If you can used them before, please give me some advice. Thanks!</p>",
      "rawMarkdown": "based on the prediction scores, it can identify which training (and testing) samples are difficult. For example, now i have: {x_i,y_i,s_i} for i=1 to N train samples. x is the image, y is the ground truth label, s is real value to indicate the difficulty. \n\nHow can i use use the value of s to improve my training?\n\nE.g. how to formulate better loss?\n\nE.g. how to do better sampling?\n\nThis is related to Self-paced Learning or Curriculum Learning. Although I have read about such methods, I have never tried them before. If you can used them before, please give me some advice. Thanks!",
      "votes": null
    },
    {
      "id": "233225",
      "postDate": "10/19/2017 16:16:30",
      "content": "<p>an example is:\n\"Curriculum Learning with Deep Convolutional Neural Networks\"\n<a href=\"http://kth.diva-portal.org/smash/get/diva2:878140/FULLTEXT01.pdf\">http://kth.diva-portal.org/smash/get/diva2:878140/FULLTEXT01.pdf</a></p>\n\n<p>Note that how we can speedup training with less samples.\n(You can always add back the removed samples in final stages of fine-tunning)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/233225/7708/easy_remove.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "an example is:\n\"Curriculum Learning with Deep Convolutional Neural Networks\"\nhttp://kth.diva-portal.org/smash/get/diva2:878140/FULLTEXT01.pdf\n\nNote that how we can speedup training with less samples.\n(You can always add back the removed samples in final stages of fine-tunning)\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/233225/7708/easy_remove.png",
      "votes": null
    },
    {
      "id": "233228",
      "postDate": "10/19/2017 16:19:09",
      "content": "<p>Another example is:\n<a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>\n\n<p>\"Hard sample loss\" is used to improved results. But details are not given.</p>\n\n<p><img src=\"http://5047-presscdn.pagely.netdna-cdn.com/wp-content/uploads/2017/10/image8.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "Another example is:\nhttp://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\n\n\"Hard sample loss\" is used to improved results. But details are not given.\n\n\n  ![enter image description here][1]\n\n\n  [1]: http://5047-presscdn.pagely.netdna-cdn.com/wp-content/uploads/2017/10/image8.png",
      "votes": null
    },
    {
      "id": "233273",
      "postDate": "10/19/2017 18:30:31",
      "content": "<p>Hi,Heng,<br>\nIn summary,the hard example use in that competition quite similar to your post here.<br>\nHere is the detail:<br>\nIn a batch,when we computing loss,we compute loss on every sample first ,in that competition,is the LogLoss of each label and sum them up. Then,I sorted the loss,and keep the Top x% examples which have relatively large loss,and sum the loss of them up and get gradient ,back-propagate... ... <br></p>\n\n<p>I  did not have a try on this competition yet. Perhaps,you can use hard-example loss function to finetune your model after 2-3 epochs.I dont know whether it will be helpful or not.Because I am a little busy and did not even have a look at the images and analyse the distribution of the categories/products .</p>",
      "rawMarkdown": "Hi,Heng,<br>\nIn summary,the hard example use in that competition quite similar to your post here.<br>\nHere is the detail:<br>\nIn a batch,when we computing loss,we compute loss on every sample first ,in that competition,is the LogLoss of each label and sum them up. Then,I sorted the loss,and keep the Top x% examples which have relatively large loss,and sum the loss of them up and get gradient ,back-propagate... ... <br>\n\nI  did not have a try on this competition yet. Perhaps,you can use hard-example loss function to finetune your model after 2-3 epochs.I dont know whether it will be helpful or not.Because I am a little busy and did not even have a look at the images and analyse the distribution of the categories/products .",
      "votes": null
    },
    {
      "id": "233373",
      "postDate": "10/20/2017 00:44:38",
      "content": "<p>thank you for the reply. i will try it out!</p>",
      "rawMarkdown": "thank you for the reply. i will try it out!",
      "votes": null
    },
    {
      "id": "233384",
      "postDate": "10/20/2017 01:51:23",
      "content": "<p>could you please share some training tricks?</p>",
      "rawMarkdown": "could you please share some training tricks?",
      "votes": null
    },
    {
      "id": "233420",
      "postDate": "10/20/2017 06:51:05",
      "content": "<p>@bestfitting In the original competition did you try taking not top x% but using tome threshold like take only those that scored loss above certain level? I know that it is hard to define what should it be, but I guess that could be deducted from the validation sample after 2-3 epochs of training.</p>",
      "rawMarkdown": "bestfitting In the original competition did you try taking not top x% but using tome threshold like take only those that scored loss above certain level? I know that it is hard to define what should it be, but I guess that could be deducted from the validation sample after 2-3 epochs of training.",
      "votes": null
    },
    {
      "id": "233470",
      "postDate": "10/20/2017 10:10:57",
      "content": "<p>Hi,Marcin,sure,we can do so.To do so we can analyse the distribution of the probs and choose a strategy.But first of all,we must analyse the model's capabilities on different samples and then do something to enhance the weak aspect of our models.<br>\nAs we discuss here is hard examples,we can also pay some attention to training set itself,in dataloader ,we can sample  images guide by the loss of last epoch(or certern iters) or distributition of categories to save time and balance the categories.<br>\nThe dataset of this competition is so large,it gives us a lot of possiblities to design all kinds of networks and use a lot of skills. <br></p>",
      "rawMarkdown": "Hi,Marcin,sure,we can do so.To do so we can analyse the distribution of the probs and choose a strategy.But first of all,we must analyse the model's capabilities on different samples and then do something to enhance the weak aspect of our models.<br>\nAs we discuss here is hard examples,we can also pay some attention to training set itself,in dataloader ,we can sample  images guide by the loss of last epoch(or certern iters) or distributition of categories to save time and balance the categories.<br>\nThe dataset of this competition is so large,it gives us a lot of possiblities to design all kinds of networks and use a lot of skills. <br>",
      "votes": null
    },
    {
      "id": "233471",
      "postDate": "10/20/2017 10:11:25",
      "content": "<p>Hi,iFighting,I get the idea from this paper,although it's about object detection:Training Region-based Object Detectors with Online Hard Example Mining(<a href=\"https://arxiv.org/abs/1604.03540\">https://arxiv.org/abs/1604.03540</a>).<br>perhaps you can search Online Hard Example Mining for more tricks on google,and pay attention to implementation details and experiments part of a paper.<br>\nSince you public LB is very good now, I think you are in right direction,so keep going and good luck.:)</p>",
      "rawMarkdown": "Hi,iFighting,I get the idea from this paper,although it's about object detection:Training Region-based Object Detectors with Online Hard Example Mining(https://arxiv.org/abs/1604.03540).<br>perhaps you can search Online Hard Example Mining for more tricks on google,and pay attention to implementation details and experiments part of a paper.<br>\nSince you public LB is very good now, I think you are in right direction,so keep going and good luck.:)",
      "votes": null
    },
    {
      "id": "233552",
      "postDate": "10/20/2017 14:21:17",
      "content": "<p>thanks for your advice</p>",
      "rawMarkdown": "thanks for your advice",
      "votes": null
    },
    {
      "id": "233564",
      "postDate": "10/20/2017 14:37:36",
      "content": "<p><a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>",
      "rawMarkdown": "http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/",
      "votes": null
    },
    {
      "id": "236426",
      "postDate": "10/27/2017 10:01:37",
      "content": "<p>Can focal loss be applied in this task？  <a href=\"https://arxiv.org/abs/1708.02002\">https://arxiv.org/abs/1708.02002</a></p>",
      "rawMarkdown": "Can focal loss be applied in this task？  https://arxiv.org/abs/1708.02002",
      "votes": null
    },
    {
      "id": "236498",
      "postDate": "10/27/2017 14:12:26",
      "content": "<p>in theory yes. but you have to be careful of your batch size. If your batch size is small, distribution is distorted. Focal Loss is already implemented in my pytorch kit, but i haven't try it yet in details.</p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>\n\n<pre><code>## https://github.com/unsky/focal-loss\n## https://github.com/sciencefans/Focal-Loss\n## https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39951\n\n#  https://raberrytv.wordpress.com/2017/07/01/pytorch-kludges-to-ensure-numerical-stability/\n#  https://github.com/pytorch/pytorch/issues/1620\n\n\nclass FocalLoss(nn.Module):\n\ndef __init__(self,gamma = 2, alpha=1.2):\n    super(FocalLoss, self).__init__()\n    self.gamma = gamma\n    self.alpha = alpha\n\n\ndef forward(self, logits, labels):\n    eps = 1e-7\n\n    # loss =  - np.power(1 - p, gamma) * np.log(p))\n    probs = F.softmax(logits)\n    probs = probs.gather(dim=1, index=labels.view(-1,1)).view(-1)\n    probs = torch.clamp(probs, min=eps, max=1-eps)\n\n    loss = -torch.pow(1-probs, self.gamma) *torch.log(probs)\n    loss = loss.mean()*self.alpha\n\n    return loss\n</code></pre>",
      "rawMarkdown": "in theory yes. but you have to be careful of your batch size. If your batch size is small, distribution is distorted. Focal Loss is already implemented in my pytorch kit, but i haven't try it yet in details.\n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n    \n    ## https://github.com/unsky/focal-loss\n    ## https://github.com/sciencefans/Focal-Loss\n    ## https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39951\n\n    #  https://raberrytv.wordpress.com/2017/07/01/pytorch-kludges-to-ensure-numerical-stability/\n    #  https://github.com/pytorch/pytorch/issues/1620\n\n\n    class FocalLoss(nn.Module):\n\n    def __init__(self,gamma = 2, alpha=1.2):\n        super(FocalLoss, self).__init__()\n        self.gamma = gamma\n        self.alpha = alpha\n\n\n    def forward(self, logits, labels):\n        eps = 1e-7\n\n        # loss =  - np.power(1 - p, gamma) * np.log(p))\n        probs = F.softmax(logits)\n        probs = probs.gather(dim=1, index=labels.view(-1,1)).view(-1)\n        probs = torch.clamp(probs, min=eps, max=1-eps)\n\n        loss = -torch.pow(1-probs, self.gamma) *torch.log(probs)\n        loss = loss.mean()*self.alpha\n\n        return loss",
      "votes": null
    },
    {
      "id": "374219",
      "postDate": "08/22/2018 16:14:28",
      "content": "<p>Hello, SeuTao. I am participating in the competition of TGS also. Could I make a friend with you. Could you tell me your qq or WeChat?</p>",
      "rawMarkdown": "Hello, SeuTao. I am participating in the competition of TGS also. Could I make a friend with you. Could you tell me your qq or WeChat?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 233225,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/19/2017 16:16:30",
      "content": "<p>an example is:\n\"Curriculum Learning with Deep Convolutional Neural Networks\"\n<a href=\"http://kth.diva-portal.org/smash/get/diva2:878140/FULLTEXT01.pdf\">http://kth.diva-portal.org/smash/get/diva2:878140/FULLTEXT01.pdf</a></p>\n\n<p>Note that how we can speedup training with less samples.\n(You can always add back the removed samples in final stages of fine-tunning)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/233225/7708/easy_remove.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 233228,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/19/2017 16:19:09",
      "content": "<p>Another example is:\n<a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>\n\n<p>\"Hard sample loss\" is used to improved results. But details are not given.</p>\n\n<p><img src=\"http://5047-presscdn.pagely.netdna-cdn.com/wp-content/uploads/2017/10/image8.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 233273,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "10/19/2017 18:30:31",
          "content": "<p>Hi,Heng,<br>\nIn summary,the hard example use in that competition quite similar to your post here.<br>\nHere is the detail:<br>\nIn a batch,when we computing loss,we compute loss on every sample first ,in that competition,is the LogLoss of each label and sum them up. Then,I sorted the loss,and keep the Top x% examples which have relatively large loss,and sum the loss of them up and get gradient ,back-propagate... ... <br></p>\n\n<p>I  did not have a try on this competition yet. Perhaps,you can use hard-example loss function to finetune your model after 2-3 epochs.I dont know whether it will be helpful or not.Because I am a little busy and did not even have a look at the images and analyse the distribution of the categories/products .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233373,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/20/2017 00:44:38",
          "content": "<p>thank you for the reply. i will try it out!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233384,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/20/2017 01:51:23",
          "content": "<p>could you please share some training tricks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233420,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "10/20/2017 06:51:05",
          "content": "<p>@bestfitting In the original competition did you try taking not top x% but using tome threshold like take only those that scored loss above certain level? I know that it is hard to define what should it be, but I guess that could be deducted from the validation sample after 2-3 epochs of training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233470,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "10/20/2017 10:10:57",
          "content": "<p>Hi,Marcin,sure,we can do so.To do so we can analyse the distribution of the probs and choose a strategy.But first of all,we must analyse the model's capabilities on different samples and then do something to enhance the weak aspect of our models.<br>\nAs we discuss here is hard examples,we can also pay some attention to training set itself,in dataloader ,we can sample  images guide by the loss of last epoch(or certern iters) or distributition of categories to save time and balance the categories.<br>\nThe dataset of this competition is so large,it gives us a lot of possiblities to design all kinds of networks and use a lot of skills. <br></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233471,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "10/20/2017 10:11:25",
          "content": "<p>Hi,iFighting,I get the idea from this paper,although it's about object detection:Training Region-based Object Detectors with Online Hard Example Mining(<a href=\"https://arxiv.org/abs/1604.03540\">https://arxiv.org/abs/1604.03540</a>).<br>perhaps you can search Online Hard Example Mining for more tricks on google,and pay attention to implementation details and experiments part of a paper.<br>\nSince you public LB is very good now, I think you are in right direction,so keep going and good luck.:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233552,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/20/2017 14:21:17",
          "content": "<p>thanks for your advice</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 233564,
      "author_name": "juhi26",
      "author_url": "",
      "post_date": "10/20/2017 14:37:36",
      "content": "<p><a href=\"http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\">http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 236426,
      "author_name": "shentao",
      "author_url": "",
      "post_date": "10/27/2017 10:01:37",
      "content": "<p>Can focal loss be applied in this task？  <a href=\"https://arxiv.org/abs/1708.02002\">https://arxiv.org/abs/1708.02002</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 236498,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/27/2017 14:12:26",
          "content": "<p>in theory yes. but you have to be careful of your batch size. If your batch size is small, distribution is distorted. Focal Loss is already implemented in my pytorch kit, but i haven't try it yet in details.</p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>\n\n<pre><code>## https://github.com/unsky/focal-loss\n## https://github.com/sciencefans/Focal-Loss\n## https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39951\n\n#  https://raberrytv.wordpress.com/2017/07/01/pytorch-kludges-to-ensure-numerical-stability/\n#  https://github.com/pytorch/pytorch/issues/1620\n\n\nclass FocalLoss(nn.Module):\n\ndef __init__(self,gamma = 2, alpha=1.2):\n    super(FocalLoss, self).__init__()\n    self.gamma = gamma\n    self.alpha = alpha\n\n\ndef forward(self, logits, labels):\n    eps = 1e-7\n\n    # loss =  - np.power(1 - p, gamma) * np.log(p))\n    probs = F.softmax(logits)\n    probs = probs.gather(dim=1, index=labels.view(-1,1)).view(-1)\n    probs = torch.clamp(probs, min=eps, max=1-eps)\n\n    loss = -torch.pow(1-probs, self.gamma) *torch.log(probs)\n    loss = loss.mean()*self.alpha\n\n    return loss\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 374219,
          "author_name": "lianandrew",
          "author_url": "",
          "post_date": "08/22/2018 16:14:28",
          "content": "<p>Hello, SeuTao. I am participating in the competition of TGS also. Could I make a friend with you. Could you tell me your qq or WeChat?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "233214": "based on the prediction scores, it can identify which training (and testing) samples are difficult. For example, now i have: {x_i,y_i,s_i} for i=1 to N train samples. x is the image, y is the ground truth label, s is real value to indicate the difficulty. \n\nHow can i use use the value of s to improve my training?\n\nE.g. how to formulate better loss?\n\nE.g. how to do better sampling?\n\nThis is related to Self-paced Learning or Curriculum Learning. Although I have read about such methods, I have never tried them before. If you can used them before, please give me some advice. Thanks!",
    "233225": "an example is:\n\"Curriculum Learning with Deep Convolutional Neural Networks\"\nhttp://kth.diva-portal.org/smash/get/diva2:878140/FULLTEXT01.pdf\n\nNote that how we can speedup training with less samples.\n(You can always add back the removed samples in final stages of fine-tunning)\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/233225/7708/easy_remove.png",
    "233228": "Another example is:\nhttp://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/\n\n\"Hard sample loss\" is used to improved results. But details are not given.\n\n\n  ![enter image description here][1]\n\n\n  [1]: http://5047-presscdn.pagely.netdna-cdn.com/wp-content/uploads/2017/10/image8.png",
    "233273": "Hi,Heng,<br>\nIn summary,the hard example use in that competition quite similar to your post here.<br>\nHere is the detail:<br>\nIn a batch,when we computing loss,we compute loss on every sample first ,in that competition,is the LogLoss of each label and sum them up. Then,I sorted the loss,and keep the Top x% examples which have relatively large loss,and sum the loss of them up and get gradient ,back-propagate... ... <br>\n\nI  did not have a try on this competition yet. Perhaps,you can use hard-example loss function to finetune your model after 2-3 epochs.I dont know whether it will be helpful or not.Because I am a little busy and did not even have a look at the images and analyse the distribution of the categories/products .",
    "233373": "thank you for the reply. i will try it out!",
    "233384": "could you please share some training tricks?",
    "233420": "bestfitting In the original competition did you try taking not top x% but using tome threshold like take only those that scored loss above certain level? I know that it is hard to define what should it be, but I guess that could be deducted from the validation sample after 2-3 epochs of training.",
    "233470": "Hi,Marcin,sure,we can do so.To do so we can analyse the distribution of the probs and choose a strategy.But first of all,we must analyse the model's capabilities on different samples and then do something to enhance the weak aspect of our models.<br>\nAs we discuss here is hard examples,we can also pay some attention to training set itself,in dataloader ,we can sample  images guide by the loss of last epoch(or certern iters) or distributition of categories to save time and balance the categories.<br>\nThe dataset of this competition is so large,it gives us a lot of possiblities to design all kinds of networks and use a lot of skills. <br>",
    "233471": "Hi,iFighting,I get the idea from this paper,although it's about object detection:Training Region-based Object Detectors with Online Hard Example Mining(https://arxiv.org/abs/1604.03540).<br>perhaps you can search Online Hard Example Mining for more tricks on google,and pay attention to implementation details and experiments part of a paper.<br>\nSince you public LB is very good now, I think you are in right direction,so keep going and good luck.:)",
    "233552": "thanks for your advice",
    "233564": "http://blog.kaggle.com/2017/10/17/planet-understanding-the-amazon-from-space-1st-place-winners-interview/",
    "236426": "Can focal loss be applied in this task？  https://arxiv.org/abs/1708.02002",
    "236498": "in theory yes. but you have to be careful of your batch size. If your batch size is small, distribution is distorted. Focal Loss is already implemented in my pytorch kit, but i haven't try it yet in details.\n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n    \n    ## https://github.com/unsky/focal-loss\n    ## https://github.com/sciencefans/Focal-Loss\n    ## https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39951\n\n    #  https://raberrytv.wordpress.com/2017/07/01/pytorch-kludges-to-ensure-numerical-stability/\n    #  https://github.com/pytorch/pytorch/issues/1620\n\n\n    class FocalLoss(nn.Module):\n\n    def __init__(self,gamma = 2, alpha=1.2):\n        super(FocalLoss, self).__init__()\n        self.gamma = gamma\n        self.alpha = alpha\n\n\n    def forward(self, logits, labels):\n        eps = 1e-7\n\n        # loss =  - np.power(1 - p, gamma) * np.log(p))\n        probs = F.softmax(logits)\n        probs = probs.gather(dim=1, index=labels.view(-1,1)).view(-1)\n        probs = torch.clamp(probs, min=eps, max=1-eps)\n\n        loss = -torch.pow(1-probs, self.gamma) *torch.log(probs)\n        loss = loss.mean()*self.alpha\n\n        return loss",
    "374219": "Hello, SeuTao. I am participating in the competition of TGS also. Could I make a friend with you. Could you tell me your qq or WeChat?"
  },
  "source": "meta"
}